Pith. sign in

Paper Citation Record · LEDGER

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

As of 21 August 2026, this Paper Citation Record lists 100 of 152 outbound references and 0 inbound Pith citation observations for arXiv:2608.06756.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06756 v1

Coverage vector

measured 100 of 152 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:34.990563Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 152 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86fa4e9e-bea2-4d0e-bafb-9151c55c7ede · outbound

This paper cites Learning transferable visual models from natural language supervision.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.427180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.427180Z digest=sha256:f05387db5a31db9c1fcfce4910ae1ab1f5dfb25d31a2fab4a7fbbed963248d9d

Observation 1b686d8d-8975-4482-80e0-a87d317f2e8d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems (NeurIPS), 2022.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.508527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.508527Z digest=sha256:0b9c53fa8d3343e4aba3a97dc6b5d58902c1ad73827c238552425321a43b7dfa

Observation 9558b71e-665b-4a39-81e4-c7be360f62ec · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.550014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.550014Z digest=sha256:2c89c1b2dcc26665207bdea7534611cd0ccc075005410fcbc80eb102e1c879b3

Observation dd610574-b780-475c-8b93-737ddcbab4ae · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.555035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.555035Z digest=sha256:393314dc7e0e0ad55b4b0b0f8a43e12696e91ea4fe360e9b4f465151783b84b8

Observation a1ce5d4f-fff2-4c7f-9659-dbf2835af667 · outbound

This paper cites Visual instruction tuning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.560003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.560003Z digest=sha256:38a06f23af12bbe99e205ffcc4ecd7acef76cbd5c48b4bb2055b8bf3789936f3

Observation 8c9e1efe-c2f4-413e-8834-fd4ec95c0b55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.564957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.564957Z digest=sha256:a72b48f008145401b5d3a4a67e98087024c42945ad1f9a4cc4d26d546495bcff

Observation 03b4107f-aab0-430a-a07c-e806c9b72ab6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.571389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.571389Z digest=sha256:f8cf4eaa568d287e5176508da6efdf398365cff43ce687950e75a36cdc8ddae1

Observation ef7ecffc-478c-4a73-9003-862bf182fa39 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.575729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.575729Z digest=sha256:132d78a6a3781dc19813388177c41e7af5a518b6f7f1e1b422aef20742247225

Observation 6128c3a3-393b-43dd-afa6-abb521a80b06 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.628607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.628607Z digest=sha256:7e8f6c19bf475c2801e450e0eaac037113270029e2b020072bacb4894af6bf81

Observation 0aefddd6-7be9-4181-b512-9e4dea5b8264 · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Code as Policies: Language Model Programs for Embodied Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.705999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.705999Z digest=sha256:2b8a0c8c7ef07ce17509cf4c348623001d2e6f11844b2b60af093465d9621306

Observation cc13936a-121e-42ec-afe3-48b8ee51cafa · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.752781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.752781Z digest=sha256:d4f9e18ab9d54cb3cd5bf84462ff3442e8a91dd6e2660fd83595aa33c8b87931

Observation 13aa88af-1dcc-432a-9a86-1f8f00e72e8d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Gemini Robotics: Bringing AI into the Physical World

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.841005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.841005Z digest=sha256:fd0d04cf38a7292ee4c2ad68ec635116d9c9cd880ee7f306c9311e98eda174f3

Observation 6bb8a24e-72ac-4260-b6c5-6e26a666a9a2 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.953272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.953272Z digest=sha256:a8beb891e1d9eff589f1fc9d6abdf0b96b25024efa5ec600deb0044a39a52dce

Observation 5666ebd1-4296-40d0-b628-97027be56e0f · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.957799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.957799Z digest=sha256:c3260fbfac64b72dd769260cc5c7fa3abc75f92b12986f066ef680f9fb5f0480

Observation d8e06185-f056-48cb-acf6-b044c6e7f7cd · outbound

This paper cites Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.962774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.962774Z digest=sha256:793a4bfe146500b6fc5ad7ec12696fbc1e4fffcae9bf8549637dc0f8e01bec79

Observation 6dbc5041-6eaf-442d-9b12-5593da0039ea · outbound

This paper cites Rynnbrain: Open embodied foundation models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Rynnbrain: Open embodied foundation models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.966568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.966568Z digest=sha256:d7cbee168334c70915800def0c5f336934b0ed1195d4ac6d4c1af62d8cb264f4

Observation 162a4f34-dbd6-4735-a7cb-c240bb768c4b · outbound

This paper cites Hy-Embodied-VLM-1.0: Efficient Physical-World Agents.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.637598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:18:32.978898Z digest=sha256:85e94c4d6ec9f8d64bcc8881483e6c06df0a6ab8a2cc71d6fafb667b1529655b

Observation f8b0690c-20b7-4e47-9f58-65448b0824d2 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Cosmos 3: Omnimodal World Models for Physical AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.991114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.991114Z digest=sha256:1103c0b46060953e758db3e764a7d7c52801f9bf6def35182f0baae8d8eed583

Observation 144621a8-978e-4ce9-85a6-2f52b3fead21 · outbound

This paper cites MiMo-Embodied: X-Embodied Foundation Model Technical Report.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MiMo-Embodied: X-Embodied Foundation Model Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.096528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.096528Z digest=sha256:14fab161ee57ae68ce7b8482f29e95267f534060c496381325c7acb85fe5a257

Observation e1a0f103-d737-4b61-a4bb-243c9b4052bf · outbound

This paper cites Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.155569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.155569Z digest=sha256:735dab1ee0cd1d0c81bd5f979145b568283e86c98ba25472b61adbca97b3b199

Observation f05fbe3c-af8b-4870-84d2-0542e86cb1fb · outbound

This paper cites Vesta: A Generalist Embodied Reasoning Model.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Vesta: A Generalist Embodied Reasoning Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.541187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:18:33.160307Z digest=sha256:8fc0961449f99479aa7e9a1558f762463df053c44ba4a0259ac5b702af30cb2a

Observation 7f6474f4-f9c4-455f-9091-189374bf72e2 · outbound

This paper cites ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.517943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:18:33.164594Z digest=sha256:1e7247155ea8660278efbdd017bfbf7a36b14178eed2c4dfb78d0f22ffb6d0c6

Observation 92858ee1-ad6a-492e-9752-bfa56571f301 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.169445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.169445Z digest=sha256:a074cab549a13230b0a289260d576eaa238029fbb97d30facb74efd57d42c98c

Observation c68e0c8e-7b94-4b0c-97be-348dbf7ff35e · outbound

This paper cites Towards embodied agentic ai: Review and classification of llm-and vlm-driven robot autonomy and interaction.arXiv preprint arXiv:2508.05294, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Towards embodied agentic ai: Review and classification of llm-and vlm-driven robot autonomy and interaction.arXiv preprint arXiv:2508.05294, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.249748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.249748Z digest=sha256:54c142690836b2ba0d26f8c315fe48ca4157c46c84bdbe507ec6a35b11c71749

Observation 33640b2c-5173-4717-a65f-a80b96a4e93f · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.Conference on Robot Learning (CoRL), 2023.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.Conference on Robot Learning (CoRL), 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.279788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.279788Z digest=sha256:bb31b932cda9fabb1cf1181e189cac93b63b6a72bbaf4f6c535fc3007adf7bb7

Observation d089a960-b2e9-4632-bf07-7fa6250fedd9 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.285071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.285071Z digest=sha256:00117981987195ddb8a7d826fdcadcd1776d8d62098d485a9ec5853b734fffb9

Observation fdf9a712-d270-4e65-9028-fbeea8192e53 · outbound

This paper cites To mix or to merge: Toward multi-domain reinforcement learning for large language models.arXiv preprint arXiv:2602.12566, 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence To mix or to merge: Toward multi-domain reinforcement learning for large language models.arXiv preprint arXiv:2602.12566, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.332402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.332402Z digest=sha256:0c67c57aeae9ffb2bdd89c0ea08bff7ac838e51d8dbf8c197795b673b811148f

Observation d008e851-84b8-4f50-9a8f-ce16372ed5d0 · outbound

This paper cites TIES-Merging: Resolving Interference When Merging Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence TIES-Merging: Resolving Interference When Merging Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.474080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.474080Z digest=sha256:7dc8ebe6f52d971f7d915d125e1f5abe174ef7b14e1a4cde77d367ec1351924c

Observation 91755aff-dbcc-4616-982a-abbdcaf6c7fe · outbound

This paper cites Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.478764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.478764Z digest=sha256:a1173a96ff3a6d263778c3efb8e1f0cfbad3899fdca876e9e3622742fcc79eab

Observation c92e48b1-89bd-47a5-9d7d-306aaa57feb6 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen3.5: Towards native multimodal agents, February 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.484070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.484070Z digest=sha256:a26fe837a0731804b7f919741f94a979e40ea2786c6a4528444d8440823dc408

Observation 4b70de19-d3b2-4d1b-a254-93af939ef572 · outbound

This paper cites Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.153623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:18:33.489323Z digest=sha256:1dc1f3ae70959f5bb17d66056b8d48e81d588b224f35d5d2aba11a7df8af1133

Observation 49450709-1715-47a5-92fb-5aa070fb0f64 · outbound

This paper cites Spatial intelligence in vision-language models: A comprehensive survey.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Spatial intelligence in vision-language models: A comprehensive survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.493556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.493556Z digest=sha256:a573c878a85127060c808c9e5b300794af1f3111dfc54a5822ca2d595cce91ce

Observation ac357ebb-62a5-4f01-a8bc-6f133836dfbb · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Scaling spatial intelligence with multimodal foundation models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.595378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.595378Z digest=sha256:84e41ca73a6169b419af79d19644782662308494d44690e8f43736329393a19b

Observation 0e8ac789-c608-44cd-9679-15df91d82e69 · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.681742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.681742Z digest=sha256:05e0eef3fa028b983f34074d79ec33751be392a8a8fe84e431802a91514dd3f8

Observation 48f5a9d4-5ba1-46ac-9242-143276a5bb77 · outbound

This paper cites Mindcube: Spatial mental modeling from limited views.arXiv preprint arXiv:2506.21458, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Mindcube: Spatial mental modeling from limited views.arXiv preprint arXiv:2506.21458, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.686415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.686415Z digest=sha256:f8216272307fb61078beaa1357348850151514b76064deb47857a87a785ef235

Observation ed101fe9-12c4-4e51-9dc3-2a8c2c769b54 · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.691153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.691153Z digest=sha256:9bd9e349d316e2726677a6cd0cd82e5a2bae329f54a9c6937ec6634ca2397615

Observation 15588c70-3e8d-42c5-b43b-e8105dab5da8 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.697230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.697230Z digest=sha256:d75ad8dc1d860c41446d21e8ae82e5a8136f74641840a1e2fe461f854277a684

Observation 52943f18-0a9c-492e-a782-7b40631aba5b · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Next-qa: Next phase of question- answering to explaining temporal actions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.742687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.742687Z digest=sha256:6fe5632d203051de9319e87b9aaaafbde7ea9bf71f38cfa3d2a6e42c604590b2

Observation 7e34d14b-8972-44f2-8160-102057ebd74e · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761, 2023.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.806130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.806130Z digest=sha256:ad64d32637dcf74b8d65dcb02068ba42a945ffd3d7bdc362090e0350dd8d0566

Observation 21664cc2-0118-4b40-a07f-dec0158dd4de · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.811024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.811024Z digest=sha256:9d25e792e50248bbe1bb34b02726fe5fe3e29c24facb6472c600eb8c1f7d7855

Observation 50b4cbe8-816b-4771-bd03-26e64f07a985 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.815978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.815978Z digest=sha256:f811eb248cce4e6ae74844bf1aec49e8acb724e4b8d7e97ef148334ab86f040a

Observation 9a915b03-7cc8-4b4c-9167-2d8e39b2fad8 · outbound

This paper cites Scaling rl to long videos.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Scaling rl to long videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.820656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.820656Z digest=sha256:82ad94f68e637c7728428b5b80c1c8f2ded9cb42252cad03fd22e93dfc31fa79

Observation 83b37b98-f0f9-4a27-b8fc-3630fc025293 · outbound

This paper cites Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.963434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.963434Z digest=sha256:83f72090daab26d120c102224de0a5c2ca1c1593439437ef6b7ae744c100c18b

Observation 2092ae7e-27ce-46de-8f73-cfb520dd1326 · outbound

This paper cites Localizing moments in video with natural language.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Localizing moments in video with natural language

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.967117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.967117Z digest=sha256:5d8a7b69d100d01a669bd93033f2ed39de52e1f1f85a86fc729c79bccb9a19a2

Observation e239d02d-f01f-4552-8eea-1acabe332c49 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Hierarchical video-moment retrieval and step-captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.035589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.035589Z digest=sha256:6f9339430f60da016cae8e9b9db1ae0f24be38d2a808a7f7c9feb49090f7cbcb

Observation eee2537f-48ac-4f0c-a942-aee7a7f2a935 · outbound

This paper cites Queryd: A video dataset with high-quality text and audio narrations.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Queryd: A video dataset with high-quality text and audio narrations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.097731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.097731Z digest=sha256:4ac448672b2a3108ae6531c622f51b68cea8afca4c6122b7efca2276ec6b7a90

Observation 4a3cba58-04d7-4264-a127-c786019291ed · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.103487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.103487Z digest=sha256:56844bbc4d0f3327cf9588ec34bca6cc0f71d77af47e9d83d19cf7a82c68ee1b

Observation 581d0e6d-922b-493d-957d-1ec240bfe891 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.108407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.108407Z digest=sha256:fd8a139a897e44fd941cf7c0bce269bbc106a942a2728876a19544b38d487e8f

Observation a9271ed1-8430-4575-b124-4d6abe3a7614 · outbound

This paper cites Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.141816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.141816Z digest=sha256:1175e7010990ccfa8afd72676ca30a28afef2ee7df9b4094376b132c65199693

Observation 5134f1df-a2ad-4d25-a522-232b9379943d · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.146624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.146624Z digest=sha256:c44b7f5fd4384b6bbecfa593f01dc519539eb599dad6440c2953ed8f105df777

Observation c77499d7-845c-43eb-8b61-5b6eb9cc16a7 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.152165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.152165Z digest=sha256:7730cd76aed111105f3189178188f747332e4401f341111205f4a0b7f4e22ee5

Observation a475d797-c9c9-496f-be8d-66ef66f0f527 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.157336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.157336Z digest=sha256:4bfd7d3352728f72cfc78179fc0b1f1d42096140824501e95ed193a3cb19c53d

Observation d7cec257-041a-42a8-b470-45aa21c02dc3 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.161372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.161372Z digest=sha256:db87766c6acefec5ed9474fbb502353e471bdc48d8cd2b395449dba9917afda4

Observation 7685dd33-d366-4b69-aa92-7fa57addc942 · outbound

This paper cites Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv preprint arXiv:2512.24653, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv preprint arXiv:2512.24653, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.165396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.165396Z digest=sha256:7e8aaf6063a7f5a804cf53fc009c11979caf54a1ce6785aac09104bdee8bd052

Observation cfeac7ad-05ee-473a-ba1f-c75afb3ad17e · outbound

This paper cites From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.169970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.169970Z digest=sha256:e1a275c8086c5cf85c6c7581b7512ce35a047401e6455749f6f4cbe0eafe2e4e

Observation ea4dbfdc-93a9-4a2e-a591-a17284b2caad · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.174958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.174958Z digest=sha256:759c0034114830dc279bf6cac4d77f17c63af4cd444fc4a4cab000884aacfe70

Observation 2b0413ca-47bc-4422-a25a-fed4ca82253d · outbound

This paper cites Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.179159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.179159Z digest=sha256:448c97dffb4730c5890b4b417d448cb22df7c2e042a1530c476403594bff9f73

Observation 28ab7762-b38e-4f95-b268-f111e381a36c · outbound

This paper cites Roboafford++: A generative ai-enhanced dataset for multimodal affordance learning in robotic manipulation and navigation.arXiv preprint arXiv:2511.12436, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Roboafford++: A generative ai-enhanced dataset for multimodal affordance learning in robotic manipulation and navigation.arXiv preprint arXiv:2511.12436, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.269327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.269327Z digest=sha256:2f4ccf83faec572ac7ef300a215eae8aae9eccb327b009bd3a77876a1cb27efb

Observation ae3bc33e-e223-41ac-bcf5-b6e32a498d83 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.348952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.348952Z digest=sha256:e04e3e9c23ed63428b5cfbe8a59047c591aca3bded4d75a3db84dd7e21b5f235

Observation 2ac471f3-3385-4990-b3f4-0f97783f9428 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence OneThinker: All-in-one Reasoning Model for Image and Video

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.360274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.360274Z digest=sha256:f7fc236703912f9fff77838a1b1219b4d779893866432fd9539878b965e4ae65

Observation 7c7f7a20-bed4-4527-abc6-eecd1df2401c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.364746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.364746Z digest=sha256:7f3eae1516721756219b9d9da33c3474db1df367a41cf1aa34e54b5163daa04e

Observation c7902afc-23c1-4703-971f-b8c6c196d616 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.371034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.371034Z digest=sha256:001c93cb6d854ae3d98948e62fa6d952d69f56bf3cd92a141b43f138f55da755

Observation da04dddd-e6f9-4246-83c3-73792ffe8c6d · outbound

This paper cites Thinking with visual primitives.Technical report, 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Thinking with visual primitives.Technical report, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.375958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.375958Z digest=sha256:0ee969020e38537113218e8904ded2a90d5cdc1c8b72d73c3a4c851f52caf8d1

Observation 16b43f36-a6d6-4399-b2f8-b37f383835c5 · outbound

This paper cites Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.422610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.422610Z digest=sha256:1d99b48d70050bd8eb6851596509eb395bfb2a5bb4dee9516c436bc0082ae575

Observation c0505013-0f33-46c6-9863-ed86be8aca6b · outbound

This paper cites Computing discrete fréchet distance.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Computing discrete fréchet distance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.467999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.467999Z digest=sha256:79ef2c383b30dc9259ad6829c9cf869af70d76cd242814004867722d048b5ed8

Observation 66a70d7f-1737-4421-a885-58eda448e264 · outbound

This paper cites Editing models with task arithmetic.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Editing models with task arithmetic

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.519320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.519320Z digest=sha256:d32a72d4baa95749d0ae341c7015826ca65fcdd50b1b9a77c9fe53089a09d515

Observation b3e6ddeb-108c-43ba-8aaa-8f31591712ce · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.522970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.522970Z digest=sha256:871de019fe9719a59cdf2407d0e4863cac814d8631822ecb898475e3dda78ab4

Observation 9240bcc3-d422-4ac5-b089-cfa9a1000620 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.527373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.527373Z digest=sha256:b97afc7858156001b04b1f9712eca998c9e6103723386bcf5fa3bfad59761da8

Observation c2b34af9-5341-4fc5-b015-cbd3944e1baa · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135, 2025

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.531761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.531761Z digest=sha256:78bdd8dad248557e9c696838e0d69b6fb8bbc869916d5cee85cd409c9f05ee63

Observation 95287c8b-3854-4d03-bffc-13c3dcd6c0b9 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.arXiv preprint arXiv:2411.16537, 2024.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.arXiv preprint arXiv:2411.16537, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.535818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.535818Z digest=sha256:d25776b393a463d1041fd362c7ea0e5fa1c3181006514e88b369ce3ac67d1867

Observation 919792d0-af6c-420c-908a-63c095ce76b0 · outbound

This paper cites Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.554776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.554776Z digest=sha256:0cc31ccdf431cea7e50a38af7e8c3622087feb8d2eb81c1eae3596914ad11895

Observation 398ecf40-eb65-45f4-905d-5072113e88ba · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Openeqa: Embodied question answering in the era of foundation models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.579247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.579247Z digest=sha256:4e08e5c24bec1e9b1926551087a1991bdcea1bce36f2d93129650def28ffbaae

Observation 4fd7c8d5-a0d7-4fa5-90c2-3f5c95f42a2b · outbound

This paper cites Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.626969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.626969Z digest=sha256:9c906aa135e73f2349f4179c2e771315e55942aac1b3ed3bd22f21169141845f

Observation 8322d1ec-0c1b-41f0-9e61-bb31e74c71f9 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.631324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.631324Z digest=sha256:0a98443f24bd4ed7886d62e63ce9980cde8d3dbe05198a3a03cc0e271bf01c30

Observation bd18a110-790c-4eee-bddb-ff6638e71253 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.635521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.635521Z digest=sha256:308e16759bc3874a337b2b517664ff1cc5a7596bfaf1a2eef4035223619bd6d6

Observation d88f57ac-3e11-4f7a-b937-26d411effd17 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.640153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.640153Z digest=sha256:83e59df76d8f68c3eb49e63b8557fa279016328ccc200317256d53556ca7cedb

Observation 9b8217a6-8b99-4e16-b871-02f201571e3e · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms.arXiv preprint arXiv:2512.14698, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Timelens: Rethinking video temporal grounding with multimodal llms.arXiv preprint arXiv:2512.14698, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.644370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.644370Z digest=sha256:171614d11617e594ea1c589145ed4d27f27386d3f9ea7c18ba2d2e382707defa

Observation 40616467-1d5b-42e1-ad5a-19fd41f99973 · outbound

This paper cites VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:36.252207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:18:34.648497Z digest=sha256:0059efafec7a59f3b0f3ff556366791971698af82fb32023dfb070429fd6ed5b

Observation 4f568cba-cfae-46ec-905e-d54e13e39fab · outbound

This paper cites Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.676111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.676111Z digest=sha256:7bdf128a2e6223c07fd48adaf5fac16d09c116ccf0b4b23be7c5a507a5c1b68f

Observation 27cbe886-8340-4e8e-af68-7b37eaf1ebf1 · outbound

This paper cites Point-it-out: Benchmarking embodied reasoning for vision language models in multi-stage visual grounding.arXiv preprint arXiv:2509.25794, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Point-it-out: Benchmarking embodied reasoning for vision language models in multi-stage visual grounding.arXiv preprint arXiv:2509.25794, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.719268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.719268Z digest=sha256:f47bb29f8225f9234f5fcb937ab34403d3b80d7029ff24720663bce636f7d6a8

Observation 71d4f895-b6d9-4a10-bc89-7c27ae3ce6bb · outbound

This paper cites Navitrace: Evaluating embodied navigation of vision- language models.arXiv preprint arXiv:2510.26909, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Navitrace: Evaluating embodied navigation of vision- language models.arXiv preprint arXiv:2510.26909, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.744188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.744188Z digest=sha256:03d272ef7549c893c4f4bc839903afb8a50d1b303b1d2071ad937186b14f1a69

Observation d7558b24-56da-4a40-869e-fc62c065f7c0 · outbound

This paper cites an unresolved cited work.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.749228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.749228Z digest=sha256:6daab9e486f21b39fa9bc9508a373025c5fe08d5a3f3d2070abf301215ff96b6

Observation c5273e05-751c-4c42-bcde-3cca43a00d34 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.753811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.753811Z digest=sha256:d96474782f10de70b72a727d6aa7873319b14409844c5f70560f49f39c08988c

Observation ca55d428-8b33-4f4a-b722-7ec2ef0e057e · outbound

This paper cites Grok-1.5 vision preview and realworldqa.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Grok-1.5 vision preview and realworldqa

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.758326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.758326Z digest=sha256:a11ceeb3813a8dc4f5ab7152196e3c2cc0d0bfae040f3028b9a3e42deaed1ac3

Observation 4e99aed2-a855-4603-85e5-e2604a61e274 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MMBench: Is Your Multi-modal Model an All-around Player?

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.762841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.762841Z digest=sha256:969ae89213335cf8defb95390dff372d9e20511766fc4644d6f2fe6ea4b422b0

Observation 7f21dbfd-f9ab-4a8e-b851-747d9f020d6c · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Instruction-Following Evaluation for Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.767697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.767697Z digest=sha256:561363d32f130f81b1f3273747a426984306f4edc83db3a27cee1b15535d2a85

Observation 0d32e7c7-a8aa-43a4-967e-aa660d459f69 · outbound

This paper cites MMLU-Pro: A more robust and challenging multi-task language understanding benchmark.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MMLU-Pro: A more robust and challenging multi-task language understanding benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.773055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.773055Z digest=sha256:7eccc2b535737418957ede7df97e3f69b8129becdfe3b4a2012562bed59d6038

Observation 085afeb7-cc38-4aa6-89b7-a271d3b69cc2 · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.777235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.777235Z digest=sha256:ff422765e94cc5b0b680b1635b2896a7abce2686bd4b365ec52435bdb974601e

Observation 4fc0a1cc-53b9-41f7-be0a-3015c08d04e7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.781389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.781389Z digest=sha256:812c79d098bafe92ce31c7ecbf96cbbcb372647bca5eaecd96cb193a6f251dd4

Observation 8f25946d-13b1-45f8-bec8-487fea92ac0d · outbound

This paper cites DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:35.961548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:18:34.803552Z digest=sha256:5c7ee55541d9b7260aa4c5d1cb21845bfa2f2d3f1aec24316cd8b20e64d58aa1

Observation 75ba6374-2fd7-4256-900b-3abf96f374ab · outbound

This paper cites Qwen3-VL Technical Report.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen3-VL Technical Report

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.858442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.858442Z digest=sha256:264e6faf098a29e620bf45017afeadb9b24f9aae518b88de1b114be5cea2adb8

Observation 40d43281-44b5-4d7c-9c6e-d0f7b39db8b8 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.918689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.918689Z digest=sha256:f241813da75e45dd85d229c24da76aa0503412af1266de065a51a81f9b3ba30b

Observation 3a4aafdf-5e88-4e14-be9a-8312c51e3e0f · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems (NeurIPS), 2022.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.954091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.954091Z digest=sha256:23e99d0f20b425c0f789652d8cd7ac0f435b66fe558177005e51241a34bcdd74

Observation 245f033b-dc99-4955-b21a-bbbd6393cd3a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence HybridFlow: A Flexible and Efficient RLHF Framework

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.958615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.958615Z digest=sha256:b785ea287781aa4d6ef96c1b0cb1ba6d5adc32b7ab9d29cd464443c47e0f1561

Observation 143c72bf-dc07-436f-8035-c98e13d32153 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.963819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.963819Z digest=sha256:1107693694c0f1e015a71aeef773ea7b72c9c2da545ad39f0f5cf12610ed4872

Observation 20da74c3-02f5-482f-a03e-a66a0670e512 · outbound

This paper cites bbox_2d": [x1, y1, x2, y2],.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence bbox_2d": [x1, y1, x2, y2],

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.969726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.969726Z digest=sha256:bc958d1cc58af7d33c8fed5d2ba4a6e06a18e7ad76867d277fea38eb869b8a82

Observation 6d543b52-9bfc-4277-9c69-ab14ab36e936 · outbound

This paper cites The left hand is holding a green lighter.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence The left hand is holding a green lighter

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.976635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.976635Z digest=sha256:ca1dabac85cb1420b3a13ce6990d049e05d4d9672558373e1971a1dfb7b1eb72

Observation 8531c9ee-2c0b-4f8a-97d7-3d689f2ad37c · outbound

This paper cites The top of the lighter is now glowing red/orange, indicating it’s hot or just used, but crucially, there is no flame coming out of it.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence The top of the lighter is now glowing red/orange, indicating it’s hot or just used, but crucially, there is no flame coming out of it

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.981728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.981728Z digest=sha256:0c236225b84025df0847cefd26a0111a24d68431a12c3a5628d992f0aadc3dc2

Observation 39cffba3-a8fb-4858-ae55-44c03298c32e · outbound

This paper cites Based on observations, is the lighter on or off?.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Based on observations, is the lighter on or off?

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.985997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.985997Z digest=sha256:32ff824ac46ab87c39a6b12fc446549e9e9faf37363fc77b8690d72842394b92

Observation e5b9aa84-f903-4067-a1e5-8c134fc50d35 · outbound

This paper cites * The club sandwich must be inside the packing box.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence * The club sandwich must be inside the packing box

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.990563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.990563Z digest=sha256:568681a0892d930ca7de9f1433ce64e80cb65dcf33ae6f8abc4823c41f449542

Pith citing papers

No inbound Pith citation observations are available.