Pith. sign in

Paper Citation Record · LEDGER

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

As of 10 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 35 inbound Pith citation observations for arXiv:2505.21432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21432 v4

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:55.651566Z

measured 111 of 111 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:09.315435Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 81007fef-4341-4caf-a768-4a4bb9bd60e8 · outbound

This paper cites Vision-language foundation models as effective robot imitators.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Vision-language foundation models as effective robot imitators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.176151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.176151Z digest=sha256:b67d0c813438bd2dfdad7c6c5920701b3cc79204b4ff098da5f5dae1cb58a8a5

Observation a679ba6e-d086-424b-9449-b56e638d267f · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.316972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.316972Z digest=sha256:21ee941ad064b4b9b73ee2c32356ac80888784e1116120803852eac9ce33b8f6

Observation fbb5c2de-20a4-4c51-9fb5-d85fe417456e · outbound

This paper cites Fastumi: A scalable and hardware-independent universal manipulation interface with dataset.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Fastumi: A scalable and hardware-independent universal manipulation interface with dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.649706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:46.445264Z digest=sha256:b12d3aa963265347f54ed26b7faa81e733bfc12215bc4ce0387481c297d530d8

Observation f7a06d05-3a29-4c1f-aa89-74229fc47b72 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.578291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.578291Z digest=sha256:27e6474446ff5b9535387816600b1395e53ab1e4e5fce6a0100dd87246cbc7b0

Observation 72bb9ad0-c3ad-4caf-92d2-0ae9f341887a · outbound

This paper cites Learning 2d invariant affordance knowledge for 3d affordance grounding.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Learning 2d invariant affordance knowledge for 3d affordance grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.457287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:46.673699Z digest=sha256:4687e9781a75d40309ff95f925c547844b176113aca76a3f38cc1142e09bc52b

Observation 523ae9b6-ae5a-4258-84ba-9c07faaa7f70 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.813552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.813552Z digest=sha256:81370a9600d21ca32d98059e8a688073579e61f03675fbfdf315d345e69c8657

Observation b16ca1ec-1d7f-44e0-94fd-da39d0af58cc · outbound

This paper cites Improving domain generalization in self-supervised monocular depth estimation via stabilized adversarial training.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Improving domain generalization in self-supervised monocular depth estimation via stabilized adversarial training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.302774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:46.961211Z digest=sha256:990c9c20640e7b412cde4319114a673f081a484f36fcb479fede3b767e4fbd53

Observation 4a1439db-2ae7-4131-a03b-830828a63ccd · outbound

This paper cites MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.040696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.040696Z digest=sha256:e338cae4291e305539fbb340632a0ccb478a245109392463eba1ce877f22d75a

Observation e4c00cb3-77d4-4214-b80e-5fa8af747019 · outbound

This paper cites ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A Star.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A Star

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.143573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.143573Z digest=sha256:cb3c2884cafdf25fa7f4cddacf7a5878faab218056487d2515056a22008346be

Observation 5f130c18-80c5-4cd6-9397-eb6338f72955 · outbound

This paper cites Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.253710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.253710Z digest=sha256:e9706283e4a499de1214b86e882b372bcd3428c27ff4f1be0f57201471e0b6b2

Observation 94eaf881-2020-449e-ab3a-2b66d3dc8480 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Gemini Robotics: Bringing AI into the Physical World

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.358067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.358067Z digest=sha256:f1ca504fc1e28661507a28a042b5dd360ad8911dff73b392b9d07cf75b80f2d2

Observation 7b973670-3451-42d6-ada0-178115d937c5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.463400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.463400Z digest=sha256:08cac31e87144f228b40496ad07e6a7d0543d10b2399c6843f7bd8a1257a3755

Observation c5d2e319-e561-438f-b073-5c14a9713627 · outbound

This paper cites Thinking, fast and slow.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Thinking, fast and slow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.579910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.579910Z digest=sha256:38b73cfcaa513213dbcfeb26cb7e3e9c958f367e0b236bb1257b4cbd49b8075d

Observation 7005a5ee-e8a7-4078-b940-90adc3152e70 · outbound

This paper cites COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.705703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.705703Z digest=sha256:9368150b4aa64cdf4e13a02d5117b7b037872ebb0e1ea3431ffbee5da54aa78f

Observation dad5dba8-74ae-4f28-932a-592062ea71e2 · outbound

This paper cites Kinematic- aware prompting for generalizable articulated object manipulation with llms.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Kinematic- aware prompting for generalizable articulated object manipulation with llms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.141533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:47.802501Z digest=sha256:a1d450788c1f0c7948b66f919142eadbdc54819592e9cbb44786b5e1f469e027

Observation 8c35e3b9-4873-4117-a821-d10f28fa2b16 · outbound

This paper cites MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.950909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.950909Z digest=sha256:44ac4e99979aae0dfe725a1180937825d06006e94802b0318b2f3c303cceddad

Observation 6dddf33d-15cc-4694-939f-d600362909b7 · outbound

This paper cites Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.084439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.084439Z digest=sha256:f06cc0f0e088d3ee05845ac4e3e748b1734d1383c795138fa317512572fae76d

Observation 597e2a01-7034-4650-b63b-32cafef1ea61 · outbound

This paper cites Robotic policy learning via human-assisted action preference optimization.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Robotic policy learning via human-assisted action preference optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.230494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.230494Z digest=sha256:c5e4d144c41f1a453763d19535d47daf233f7bf91986d68dbacb3c55bed4d65f

Observation 92e6060e-9b8b-44a7-a2c4-33d3bfb0e455 · outbound

This paper cites Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.991650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:48.361825Z digest=sha256:6378587082554794742a6be9c4312b23252b40fc5cd281c0a53f76b958ecf83e

Observation 85cd5166-56d9-467d-ae2d-9aa866a98f5e · outbound

This paper cites Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.487515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.487515Z digest=sha256:d2dc79392219f6c5bd1194b022b496f990a0ef7e814e600af308bcc905227c9d

Observation 1aa81f56-cba4-40c7-b47d-aabcf0641517 · outbound

This paper cites Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.854828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:48.605623Z digest=sha256:297c13cf5b77669088fe01cb0aab65688703a66b75a2fe11fdb8bbdf069ef9ef

Observation 1af8d243-7c3a-4325-b55d-d0415085ca79 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.688344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:48.729701Z digest=sha256:f38b3e3c0e1c68cf2a884a94ddd2313dbaead7e1914c4d25ae0e1d3fcc0759d0

Observation 33ca2d67-8cba-48d1-8c09-e441a74ddea4 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.892757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.892757Z digest=sha256:565e56701e4c0385c73e83da527d3010d1eac348547605427c5935d1fee6a010

Observation afbe1c2c-f490-4a80-a099-310c6b6a865f · outbound

This paper cites A dual process vla: Efficient robotic manipulation leveraging vlm.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model A dual process vla: Efficient robotic manipulation leveraging vlm

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.459634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:49.081807Z digest=sha256:27780c849bf49ca8cd8a4b261721e0f01eafcb4357edd9a8e6783706066f12b8

Observation 0593f14e-de49-4900-97f9-ced694cb1150 · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.257537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.257537Z digest=sha256:c22dce61507b842ba9701d96ee20018b09239f729a61f210683d0f76320c6fee

Observation fbc7f4d3-eef8-4e1c-9eb2-c8f2cbfe781d · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.385522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.385522Z digest=sha256:ce6e1f7dca6415ab73830073638832a4085b81567174db98e12c2a806126df24

Observation 5ca8f132-157a-4cc0-a0b7-b44fbf22ab28 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.511193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.511193Z digest=sha256:351238bd211c3b77e15c2d8ea3dce55827d7d1509b8ae69caa23ae78e266a66e

Observation 3c1db6a1-da02-49f4-a743-2c29061bc1a0 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.614480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.614480Z digest=sha256:abf0665746b3d28c30fb74c5fdefdf29eefed0359a86a25af77af24183525384

Observation faf03009-fb2c-461a-bd4d-14ef462cf812 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.760140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.760140Z digest=sha256:6d1cb9d3601406016c5bc9f15caea957753e71849ccff319fe6c16b3491d9f8d

Observation f3b374f1-6e31-40ef-983a-45e2b9a5fa7a · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control, 2025.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Helix: A vision-language-action model for generalist humanoid control, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.206921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:49.915193Z digest=sha256:d8d6f97d19d94c186f142fada754c311bc7e499b032f57c3487ebf9e16425462

Observation 32fb4798-d4e5-4b97-82c7-86ae6160cf74 · outbound

This paper cites Flow Matching for Generative Modeling.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.042676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.042676Z digest=sha256:d9ef1e99302a96a4abd95b64ab93815f4503fd4d5dbfa6194e27fe25284b21ca

Observation e36e8996-aef0-42d3-9da8-5e096f4a255c · outbound

This paper cites Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.178954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.178954Z digest=sha256:bc67c2497bfef7a2da42e3011676651c34b0509e853e5068d2b2f4ae51401da3

Observation 4e2d2e85-f720-40dd-9f4e-71daf55ff675 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.325026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.325026Z digest=sha256:c04abe3f6bd910e74d7ad67643ebeaa01f4f01e6f86f8db28f536fbad77065f2

Observation ed099548-4729-4806-a43b-4047d2b92ee5 · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Evaluating real-world robot manipulation policies in simulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.079764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:50.473479Z digest=sha256:2434c30425a4484bb6f24b148745dd39ec0d0fbf1854959b6acb16ec24e7438a

Observation b6980a7b-23da-4fac-9da9-898329f2d91a · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.601029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.601029Z digest=sha256:f45a3f4987b57c3adf5d247cb0693609319dd40d634b37461f6112e1dfb1ccb5

Observation 4661f28f-b4fd-428c-8cab-dd0ea2467d6c · outbound

This paper cites FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.714630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.714630Z digest=sha256:d753dc20a7efddf672fda7d4d396f3c95a33a866dff6585bcc80d60675c24e7c

Observation 257b1298-91e4-4d6a-b99e-c8e12cacef87 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.849364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.849364Z digest=sha256:16a2b5b021fd2db5294eff5184728a5a1a2b57775d159fb67cef2aef6eb0b525

Observation 15194271-763a-4a9e-ae20-58bc0171772c · outbound

This paper cites RoboMP2: A robotic multimodal perception-planning framework with multimodal large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RoboMP2: A robotic multimodal perception-planning framework with multimodal large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.951725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:50.961129Z digest=sha256:c2fc9f19efd538d5561f2421a89fca8646eeee5ec216ed96849b74f59f3b7ee0

Observation 6f71dcb0-3c70-43a6-b824-801f1038533a · outbound

This paper cites Universal Actions for Enhanced Embodied Foundation Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Universal Actions for Enhanced Embodied Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.178608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.178608Z digest=sha256:c538f417509799b3788aa3c9ba2485345c6170b1160fdd982b45e6011dc8bc06

Observation 98106cf3-bab8-479b-bef4-aedf443dc267 · outbound

This paper cites Learning causality-inspired representation consistency for video anomaly detection.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Learning causality-inspired representation consistency for video anomaly detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.715736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:51.301356Z digest=sha256:6fac604c3b754fd2fa5a8e1312f5c4f2ce6a623f74a6bfd395c8945aa235ecf1

Observation 7987215a-b994-45c6-abc3-f9d0c0442d99 · outbound

This paper cites Pali-x: On scaling up a multilingual vision and language model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Pali-x: On scaling up a multilingual vision and language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.525438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:51.416783Z digest=sha256:a336738481c0590060ef9039e1fb5a7bb2a2d4fb306cdd3637cff9d4f5a10a99

Observation a686961f-ed9d-4b19-866a-bda44f39afa6 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.312277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:51.548815Z digest=sha256:fec9619d99a8e19510d51e82d6169b33a1b9b0b978cd857cadd493131f4aca41

Observation f30ddaf8-0366-41b9-ac3c-a4e833fb91c2 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Open x-embodiment: Robotic learning datasets and rt-x models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.120220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:51.664459Z digest=sha256:71af3c3e70bee6b41b754bdf9326d6028b114099dc954bf6526d2a4368f88aa2

Observation bb30c695-ea59-47f2-b80d-691d8ebc69ca · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Reflexion: Language agents with verbal reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.772547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.772547Z digest=sha256:a81bd7fb63d3b81dac9d5a77e45a46763a877da315497d2efb412dfdcc29af40

Observation d0c08af9-ddd5-4a01-aed7-1077f92b290c · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Reasoning with Language Model is Planning with World Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.884717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.884717Z digest=sha256:198567ae06dc5b2301eea0e5bca0ba52b8160cbc49c0250e06733ceeebf0c082

Observation 76800bd8-a57b-47d3-be37-85e5bc17e353 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.985432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.985432Z digest=sha256:d010f8224833b29bd630ff5c9da3f8b001ecc5d1b05639c9c9d32cad404938ad

Observation 232a0add-3c70-40f2-85a8-dc5ba1c1ae8a · outbound

This paper cites Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.124767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.124767Z digest=sha256:9ba7d2d86e9d688aa9b22c0c4191dbc282a3918b68c737726353be1e13d7d209

Observation c515e4aa-907b-4748-b0fe-ab70894abfb7 · outbound

This paper cites Mllmguard: A multi- dimensional safety evaluation suite for multimodal large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Mllmguard: A multi- dimensional safety evaluation suite for multimodal large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.909259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:52.282259Z digest=sha256:e25d3884413871d208d0bc576cdb2876608acae7a87094f8e9a96f94632a3e40

Observation 51c50869-8ef4-42d7-9746-4530cff6eec3 · outbound

This paper cites SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.387699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.387699Z digest=sha256:0069b323b397739ae55c0ccf2d96ae732606a2ef51380e9babd1fb0f6097b239

Observation b14d58a8-8cfa-42de-b0fc-c8e86a6fd48f · outbound

This paper cites MorphMark: Flexible Adaptive Watermarking for Large Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MorphMark: Flexible Adaptive Watermarking for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.513946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.513946Z digest=sha256:9a49d3c09099e1391c11b0811f4b62c43bacde6a2e964ecd791a0ebc0dcfe198

Observation 0ade5a7e-5cda-44de-b5bf-de1ce2217f55 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Tree of thoughts: Deliberate problem solving with large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.729436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:52.622187Z digest=sha256:19c258632a0cd2b11b9fdf1db039c6576b83f14ab222ea829b1f683d35882070

Observation 52dd79ae-1572-46ca-b6ee-a04dcf031ad7 · outbound

This paper cites AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.754745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.754745Z digest=sha256:7914e8d708740addd2932b17edbb92681f81d94e6c29f798a55831288ea66ce9

Observation 56722f56-4188-4b25-aede-774f1ca05879 · outbound

This paper cites Sets: Leveraging self-verification and self-correction for improved test-time scaling.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Sets: Leveraging self-verification and self-correction for improved test-time scaling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.896307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.896307Z digest=sha256:2466eb03d9fe69028431abdc9a5aa17e39b34482af3d684edc4061db04cf4569

Observation 23c7015d-acb4-4e75-9d04-2bd052aeccba · outbound

This paper cites Interpretable Contrastive Monte Carlo Tree Search Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Interpretable Contrastive Monte Carlo Tree Search Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.995028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.995028Z digest=sha256:9702dc50ffc475d5a7fc18c56994aadc098a6ddcf575c10745fe5c6117b77a47

Observation 51468219-21ee-4e85-a594-c758e0825966 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.140367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.140367Z digest=sha256:8e28dcbc5842d9752b27fa26a36190c2d23e27e70ff12b223c44484a15308345

Observation b0acd536-3231-4720-905d-cf6adbc16476 · outbound

This paper cites Cascaded Diffusion Models for High Fidelity Image Generation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cascaded Diffusion Models for High Fidelity Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.244353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.244353Z digest=sha256:643a1eeb11d33dd84caacd523e8678f54635c733ee4cb3ced367e43fbc8b9285

Observation 55a1e608-4ed0-4fa6-bde0-4ddcb2a9a26a · outbound

This paper cites Revis- iting multi-agent world modeling from a diffusion-inspired perspective.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Revis- iting multi-agent world modeling from a diffusion-inspired perspective

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.327795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.327795Z digest=sha256:4656d846a4c1c7f75f102094b3b58c57ce93ff3f0b1fede37a909a89476933bc

Observation f73061fa-010c-4826-9dc2-865bbd20f8fd · outbound

This paper cites f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.228091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:53.420181Z digest=sha256:cf26a150ef64fdb2de88563ebd133d8a55205759e561a17b4a5671da27a0c57c

Observation d2f7f3f1-10d6-4d08-ab54-bf8ea252f4b1 · outbound

This paper cites Bring Metric Functions into Diffusion Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Bring Metric Functions into Diffusion Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.005341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:53.554629Z digest=sha256:1f46985acc0bdc0582f6d7727f78b5cd564cf8c571adcb6fb9ce3951d3299973

Observation 50763b71-7e9d-4562-b70a-c7dee4706c6c · outbound

This paper cites Spectral-cascaded diffusion model for remote sensing image spectral super-resolution.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Spectral-cascaded diffusion model for remote sensing image spectral super-resolution

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.526003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:53.717819Z digest=sha256:ae4d4b9d1c98f09240be07454232b6a6390257f5598658df65587eada29158af

Observation f51f89c3-d1cb-4027-84ce-d0d70f442a8f · outbound

This paper cites High-resolution frame interpolation with patch-based cascaded diffusion.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model High-resolution frame interpolation with patch-based cascaded diffusion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.305583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:53.835285Z digest=sha256:18a953774c1a0f08d85989f505dc54587432c83fc95779f96bf87ec326faabb5

Observation c527b747-cc31-4b73-8b80-f6f1e271296e · outbound

This paper cites Cascaded diffusion models for virtual try-on: Improving control and resolution.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cascaded diffusion models for virtual try-on: Improving control and resolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.127619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:53.990234Z digest=sha256:263ee4e9b47f82eea88c901e0a2970f2dd973357c64041adb462de3f3553a8d0

Observation 2c28b4f4-4442-4192-947e-bf9bac79cfd2 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.126278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.126278Z digest=sha256:f04f7ebaf67f4ea3b3ca769f1406eca6abe83e314876d0f6ee2ec2329511f254

Observation 05da854f-a948-4814-ad54-9ffdfb1cc9e1 · outbound

This paper cites Pre-training for robots: Offline rl enables learning new tasks from a handful of trials.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Pre-training for robots: Offline rl enables learning new tasks from a handful of trials

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.971727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:54.254443Z digest=sha256:a8ac65d346227fffc20b1aba078dbe8bc2b1d5a84436fcfe941f77bd3566421e

Observation 287d1e5a-a6e8-4633-a58d-c5f73ef9088f · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.814329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:54.384981Z digest=sha256:22b8e88c3078d6bda356a775d14d5f21cc57a6fcf453f7541502d417e7e3235b

Observation f5826893-a8b4-4a18-848d-bfba099c12f1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.509544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.509544Z digest=sha256:a638e0dcbe5a92ff9b203a57f08373d28343c2587674fd874cf1d61d39b8c01d

Observation b73c3ab7-a074-4ed8-bac6-3ee1f980e446 · outbound

This paper cites Octo: An open-source generalist robot policy.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Octo: An open-source generalist robot policy

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.674777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.674777Z digest=sha256:4ba700f23b5cddfe0e5f090c8abd53079f3bf5c68de82d49aec66bdc007526a7

Observation 949e21aa-1443-418a-a674-d4b10016708d · outbound

This paper cites Scaling proprioceptive-visual learning with het- erogeneous pre-trained transformers.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Scaling proprioceptive-visual learning with het- erogeneous pre-trained transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.645007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:54.796683Z digest=sha256:8daf40397c0bb0cebae48c375ae5bbb22c41490f8111e4f2f52c80e3280c4f6e

Observation fb8bd160-8b34-48c3-9899-1b76a00db3e0 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.889344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.889344Z digest=sha256:8778d4c3f1d828c64346f0f6fe4a2e5d1bacf464dba57395b2f67392f04acbf5

Observation 4ee513b6-aba5-475c-9d35-14fb20a3afc8 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.004692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.004692Z digest=sha256:a6c441debe154c824eac2441f2a002253082df3bcb40226b3d5a37719017801e

Observation ad6b665d-616d-44f5-bdca-262e55c9930f · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Diffusion policy: Visuomotor policy learning via action diffusion

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.157444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.157444Z digest=sha256:8e095be447d9756298c6a346b95b1d8d34798bc2315b8cbe7df71684fc2b596f

Observation 7ac1ebce-6d32-4038-8412-5bb69cd72bc6 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.223798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.223798Z digest=sha256:f7ad81ac506bdef712f2868c057343ddad05756574852ef420328211685c5f7b

Observation ae1bf471-dc67-454c-9dff-58f2d586f78c · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.329461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.329461Z digest=sha256:e3107d8f8d7d6815bb380e8f03b23b4fe2661ac9f6ec0ffdae4bf6ab0565267f

Observation 1791c3be-837b-4f62-b6bf-1553d4807b86 · outbound

This paper cites Continuous control with deep reinforcement learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Continuous control with deep reinforcement learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.428373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.428373Z digest=sha256:a7e00c2eabab6de1932723591599b81cefa9cc777f62615dd90b78a361b50e03

Observation 9b18831f-3bec-46ab-82ca-e17095e176ce · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Addressing function approximation error in actor-critic methods

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.417704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:33:55.564589Z digest=sha256:a64cab325f469cb5a25c38946935f674d7249a0bd2ea84771e4d9937efcdfbf4

Observation 5d8e87df-c426-4323-8554-e9c8d0a2fea1 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Soft Actor-Critic Algorithms and Applications

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.651566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.651566Z digest=sha256:ea7fdeda80152389c1990c5a471085ef5a87344717822fcfa6498a3bb328cdba

Pith citing papers

Observation 7f957329-1602-4060-b1fb-d8301b5ef22a · inbound

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving cites this paper.

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:19:42.874884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:19:42.748573Z digest=sha256:51d2c39b2d5d9d66a2c757c5e5881530042af760ae0e19b9cddc99a9da966158

Observation 9943f6b5-3f45-44ca-a277-c0385c11385e · inbound

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface cites this paper.

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.315435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.315435Z digest=sha256:d50520d12ccc79bcf41014f3949863c1a33292b6917579bdcf37e9fa390eeef1

Observation bf1e6f74-a5fa-4853-a89d-9d6deb14871f · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.581802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:ca149ae8daced033b4a093101b19d98cbc88fd1595ed358624452475659aad26

Observation e6c42ed0-3a1e-41ec-be9f-2ac5e9d0b556 · inbound

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions cites this paper.

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:42:47.069003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:42:46.978148Z digest=sha256:8ec12f4d89cb62cb49bf8fe70afd9b54c05e0c41b041016f8650f27dcdecee97

Observation 3edc250c-f891-48c5-987d-ad183462ea08 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:39.902748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:5f4f17aaf9447250671531e760036b4c99748798b9e3b9063208417ac872e0e9

Observation c5a28893-c809-42c3-bf40-11c38ea3f232 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:11.780961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:11.780961Z digest=sha256:a02491bd38b3d1998b30d6233d352f5d23a19e2d1bf7bbcc99b81bbf8ddb7808

Observation fa07afbc-06fa-432a-b218-afd5a0eb7310 · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:26.042642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:a4a625d4df094c1b58f58501f958751ab09b2ebca225179d965aaa0daf2cffa2

Observation 150d728d-7652-4a3c-bc24-46f9e1ba84ac · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.030889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.030889Z digest=sha256:2c4b464c6785b08f593c3899df4451a38fc0116fd9d8abed7739d041d4941563

Observation 73a9281f-48a8-4fa0-a5d9-b79e4083dc45 · inbound

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data cites this paper.

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T12:10:54.741769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:10:54.741769Z digest=sha256:eeff6550b2024d5a6e8a1dc65e5203806a74528b59e779d1281c80f8773670a2

Observation e773579a-6a7d-4e0d-8a2f-5b3e8483c3fb · inbound

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models cites this paper.

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.167968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:50:02.670780Z digest=sha256:6c2ed1fdec13e4ff1e76cde514ec1cdfeb80dec0b879d6afdbf6515716b8e041

Observation 8deee7bd-f6c2-4f84-b333-13298a69cef4 · inbound

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving cites this paper.

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.527563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T11:56:40.836234Z digest=sha256:4db7ca720630a8be23d081c883fef0ad12107dcb99a1cec90a17934ac281326f

Observation 635e0017-9b9f-4d17-b2f7-0e961982ed91 · inbound

Spatial navigation in preclinical Alzheimer's disease: A review cites this paper.

Spatial navigation in preclinical Alzheimer's disease: A review Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T19:53:53.219607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:53:53.219607Z digest=sha256:dc09a538d0034be78a184f81db3a0ae7f3112b6eb197e1bcf42e78e39f929126

Observation 5241d33f-930f-4b74-8528-f7da22a0e8d4 · inbound

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models cites this paper.

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:43:19.018696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:40:30.771335Z digest=sha256:36edfb740a66cb421467b8976435aa76060f66afdc8f8293e7cce120881257c2

Observation 90f67b5f-aacf-4d18-a407-f2dfa1a89af7 · inbound

Deep Image Clustering Based on Curriculum Learning and Density Information cites this paper.

Deep Image Clustering Based on Curriculum Learning and Density Information Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:03:28.527131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:03:05.442554Z digest=sha256:a9ff0e16b6b4d5547afdbc0f25efa526d4d96c584351cf819c5c6bc8d98f96ee

Observation 28a5c88f-922d-4a9e-afe3-e79a16620ddd · inbound

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models cites this paper.

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.503826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T16:57:52.595165Z digest=sha256:060ec5feac27b95f5b2e19c00559b12a3789069ef890c73e565915e59cc9eaae

Observation bce58944-e575-4e19-aeeb-2f96c4ab1f0d · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:04.809662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:14:03.865035Z digest=sha256:fbe23b143a2c19d6e50ff32821fda85da1f68e095a78998e55f0f222e0b29ec4

Observation c66e8685-deec-4791-a0d1-cde96d09ed9f · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.064730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:57a4388cffd44890b601bd67b8a0d47f26f2ff95b0a340ec32c377a54596d455

Observation c8b0042f-638e-4de8-b0bf-1858489ed884 · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.358813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:596e4d043dcec2ccd5e4ea9e5913bc783c825a46ec8d72415fad8ada1b6d3470

Observation 9237db13-56b0-4f9c-8a7b-61c5ff07d408 · inbound

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control cites this paper.

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:30.361693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:47:16.205357Z digest=sha256:f1575f6dae894d69fcd937e3dc7b7533c8742cabd822e192138c3eec9e7f9b07

Observation 2913645b-6fa2-4e10-8c4d-453eb1512fbb · inbound

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control cites this paper.

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.759639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:18:14.307099Z digest=sha256:560cf1f590b42d35d0adf5e86c2f1906f10d2389b1cf1969c51394f5d3c1668e

Observation 30972e20-377c-40dd-b05c-ff820df2d8c4 · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:12:18.267361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:94264def516bc5964fa81fdf277ce58d17fd5727891899d9bdf59bf0a8cce2ea

Observation 0c5495cd-9e10-45be-8883-67a3eef00e56 · inbound

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA cites this paper.

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.737640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:57:25.010339Z digest=sha256:d43a3976de5bb3506d09ce6fd9622fa7d2ab1f6fb14286e9af432608afd2bf8d

Observation de9b08b3-5e48-45e1-9deb-259b4f49c5da · inbound

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA cites this paper.

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:13:46.004799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:13:46.004799Z digest=sha256:9b327883c7928cff2b52403a5f4f0aa82958886df6f36cf4c94e4bd7ee297b0d

Observation 5e3fd18b-156e-424d-98fa-ba4770f928df · inbound

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space cites this paper.

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:58.507228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T00:57:58.312567Z digest=sha256:0bcdce025d56c8e24f9d5ba2e6c114a6022d668b9c90b4262ae1ee3e873ae92e

Observation 53e4a2ea-4c7a-4e32-b782-a1ccb57258a1 · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.647699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:b87259f9005c52b728d89efb188e66c79e70bdcdb6a6bfc7ff898fa8af9db596

Observation 8cf1726e-96e8-4ac7-8326-8919eb2502a4 · inbound

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation cites this paper.

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-27T04:40:32.814702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T19:29:00.117285Z digest=sha256:8b8578af56cdae1a1a93084268156ad26a6568c0dac57bb94f85e0af9b133f61

Observation a96ac403-2fb8-43bc-bf7f-951da40ff893 · inbound

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation cites this paper.

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.756964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T04:48:28.868562Z digest=sha256:9ffd0bcb39ddb028c8dc135562f48cc25223180686a0d4c85154f94db5b346dd

Observation 1c1c081c-dc97-461f-81d4-34b61c42e28c · inbound

Recursive Self-Evolving Agents via Held-Out Selection cites this paper.

Recursive Self-Evolving Agents via Held-Out Selection Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.572040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T11:13:15.923479Z digest=sha256:c8a2f359bcac8b0c1db9a54eeda29cf9ca6ddeedad610875304dba00bf8e5e1a

Observation 6d0ce5b9-3a94-40f1-b56c-203336e675e1 · inbound

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning cites this paper.

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.827744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:56:45.197407Z digest=sha256:f8f48d0230b218b8148c267693cd713da5d24f37a5d76cd16828a76f0d496f0a

Observation 984e3d47-f327-4463-81e0-474103ed5288 · inbound

ROSA: A Robotics Foundation Model Serving System for Robot Factories cites this paper.

ROSA: A Robotics Foundation Model Serving System for Robot Factories Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:16:52.678721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T11:11:39.584755Z digest=sha256:6d1f28a9f44d05c776ce83352a47362c7f3c607168faa8c1a1251dc7a7fb09a9

Observation 67504498-1254-428f-91dc-015b7465a512 · inbound

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models cites this paper.

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T00:12:42.173815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:12:42.173815Z digest=sha256:9b7c3498cc6130e28e1452998356b5ee51cfd2d79b149a87633c4ab5344f59a6

Observation 73f704b0-a5e4-469e-b3e9-0953aa79b3a5 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T12:10:21.115628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:10:21.115628Z digest=sha256:ed51e1fb6ca23a328fda887f352c1d0863aa6e47b0b12ae59a68178b31b94527

Observation 5efd68fd-27c8-462e-ab83-5827f5d3a342 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.336447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.336447Z digest=sha256:5f5a8fa5d6e3200df0b46481a9ce32d495112649c49536aa5d7c0d0c362a4fb8

Observation 9c83eae8-90ec-454b-8d38-30cba65811db · inbound

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking cites this paper.

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:30:57.030649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:30:57.030649Z digest=sha256:c7150bc66948a3a64939eaf05b806e2f9d55d35899bcb6aee0a6723a5a7e7176

Observation d8f48e30-f03c-4298-aa2d-b3bd694feb13 · inbound

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation cites this paper.

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T19:56:57.191876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:56:57.191876Z digest=sha256:3f379773edfb4a3f7cf314f2484533dac714aae75dd1c4ad92580ea3f37b421b