Pith. sign in

Paper Citation Record · LEDGER

Universal Visuo-Tactile Video Understanding for Embodied Interaction

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 6 inbound Pith citation observations for arXiv:2505.22566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22566 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:26.504997Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:47:20.035585Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:05:40.612441Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43cfddc7-1d6d-41b0-af88-f4bf9e6f1ad5 · outbound

This paper cites A review of tactile information: Perception and action through touch.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A review of tactile information: Perception and action through touch

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.774596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:19.560774Z digest=sha256:bf2b009eeb59eda576fd5b889a064fe870157fa5c9428674a88a904c7efd1325

Observation fbfc6331-cc9d-439c-a6d0-4d4eab314da0 · outbound

This paper cites Task and material properties interac- tively affect softness explorations along different dimensions.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Task and material properties interac- tively affect softness explorations along different dimensions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.547098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:19.672847Z digest=sha256:bd931e4aa12280d3e0f23f765a26c710fe32400bd3e144a5233cd9661bf635aa

Observation 8a8dadab-e645-4958-b1e7-f92336479787 · outbound

This paper cites Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.284124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:19.825495Z digest=sha256:883675febafb0e725b3bc3c529a6e78999a5158fb8ce46b98334504e98824730

Observation c4f5cbb6-2d5b-4f5f-8785-90ec4adb4fd0 · outbound

This paper cites Qwen Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:19.929844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:19.929844Z digest=sha256:bfe73bae4562802463411cda3fb53db4ecc05e5f6d9f621b32a3f707402c1978

Observation 0301aa58-7e92-46ab-8142-1e633397736b · outbound

This paper cites Qwen2.5 Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.070560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.070560Z digest=sha256:4efcdfcba8ebcd465221a568c25f74d8c6f5c91f20937e69a4ecd1f2ca99af25

Observation 2d442cf6-3280-4c99-8959-021be48de50b · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction High- resolution image synthesis with latent diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.167450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.167450Z digest=sha256:e021a517530007c95ae09c4e84c912bd361747f05b7f0ce574a42ef32cff4a9e

Observation 98429ef1-db62-4cff-bb88-2344d067e4cb · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.291420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.291420Z digest=sha256:0cb15ab1187cf4296e204268f4c7d48b5ccd0d3de7fefae574b639ca21586bd2

Observation e1106966-9d92-452e-9005-3dcd31fb4d2c · outbound

This paper cites CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis.

Universal Visuo-Tactile Video Understanding for Embodied Interaction CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:27.119977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:20.431420Z digest=sha256:c1fa0be8cc84137df2bf88d9a4f73ce917d0c5a0647dc14998c187b2c245e0f9

Observation c811a921-8343-401f-ae5f-591f7be8ffad · outbound

This paper cites When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective.

Universal Visuo-Tactile Video Understanding for Embodied Interaction When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.073625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:20.545737Z digest=sha256:fdf4df88d6725a74047bb0097f05e3edbcd7d590145537186fc0e91037b76186

Observation 1a6ffca7-6dd9-4dc7-9421-2c0cd24964fb · outbound

This paper cites Gelsight: High-resolution robot tactile sensors for estimating geometry and force.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gelsight: High-resolution robot tactile sensors for estimating geometry and force

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.833076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:20.676655Z digest=sha256:2c7d04b95985ba0aad5b5c696f8195044b1b79bc37e2fc8a2a5757d237f2368b

Observation dfea851b-8f91-4e67-a99d-ace2f346f1a7 · outbound

This paper cites Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.535134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:20.786246Z digest=sha256:f7432825d0e25a3494613af7cc5759b8b71ae07e1f3745c982f10dfb1de58255

Observation acebc66d-26bc-4d24-abc6-ef6bd4c66b39 · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.888419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.888419Z digest=sha256:0aac0af7e711fe4b78cbe0ae671adb340d879716f9e7ca24e39cab80b7236894

Observation 76ef809c-43e9-4f30-8cf1-93e8632aa829 · outbound

This paper cites Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.238115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:21.005344Z digest=sha256:8ce395d1aed8016c8da9e41f4ff535693b9e8968c074f10ab0d0ed1911e6d12d

Observation 624d1181-696b-404c-ab7e-f6392e065556 · outbound

This paper cites Transferable tactile transformers for representation learning across diverse sensors and tasks.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Transferable tactile transformers for representation learning across diverse sensors and tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.979642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:21.115468Z digest=sha256:b892ac29320d8a1a2ad18c485fc5a7a1b6cc876e944b90a0de8a5a8bec5c8166

Observation 793c2a97-ea4c-4375-88da-a04937c09976 · outbound

This paper cites Octopi: Object Property Reasoning with Large Tactile-Language Models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Octopi: Object Property Reasoning with Large Tactile-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.199940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.199940Z digest=sha256:f81dbed09c99948d066fd0918e5344ab39c41c4cdacb2d8762436d270e573018

Observation cdda796b-c152-44fa-914f-0ee8ff653e78 · outbound

This paper cites A touch, vision, and language dataset for multimodal alignment.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A touch, vision, and language dataset for multimodal alignment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.655885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:21.303672Z digest=sha256:ecb34ce9a0b65974cf9a2143de9c7e42994a514e0d1191f163287b15711ca3df

Observation f7fbd7ce-31e4-40ec-bb6b-37ec7f3bed1b · outbound

This paper cites Binding touch to everything: Learn- ing unified multimodal tactile representations.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Binding touch to everything: Learn- ing unified multimodal tactile representations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.437900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.437900Z digest=sha256:71c3efb2c2a4008e0ddc90d775c19faa8890d9bbc96b47ccbe1e454b83d326e6

Observation 7ea52084-156f-4997-a910-f59c5759fb43 · outbound

This paper cites Sparsh: Self-supervised touch representations for vision-based tactile sensing.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Sparsh: Self-supervised touch representations for vision-based tactile sensing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.436273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:21.550534Z digest=sha256:947fcfc515c712a1903641b4a8098f9577494579308c76f1a7b1ea48837b89a6

Observation ebbfd123-0d6e-4b15-a32c-6d0014f2f269 · outbound

This paper cites Visuo- tactile affordances for cloth manipulation with local control.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Visuo- tactile affordances for cloth manipulation with local control

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.215381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:21.673000Z digest=sha256:ec9447b087711359fd8aca8ebdd4c73c492db8fae40347d72a74d59a63ae326a

Observation cbec3edb-3dd7-4112-afbc-bf4784575848 · outbound

This paper cites A Survey of Embodied Learning for Object-Centric Robotic Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A Survey of Embodied Learning for Object-Centric Robotic Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.815777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.815777Z digest=sha256:a15c2bc73e6c5b0ff430af00030296e502c4a7a9cef0910ce42ef46dbc3c8baf

Observation 208acdd1-c51d-49ff-aee7-c0666b5a9c53 · outbound

This paper cites Touch and go: learning from human-collected vision and touch.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch and go: learning from human-collected vision and touch

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.953063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:21.941510Z digest=sha256:a2e5e0a715a2623cf219276e383041be1a32e04be3755dd5f075cf505759f4bd

Observation 392610d5-dcb9-4574-992d-ee451a827144 · outbound

This paper cites Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.696731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:22.066880Z digest=sha256:d9fd46e793207819cf41031c1070518845b2e2183b9c717e8c60c26c033f1a72

Observation 0d1a6651-3972-4388-a3d7-4f43fee1f54d · outbound

This paper cites Objectfolder 2.0: A multisensory object dataset for sim2real transfer.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder 2.0: A multisensory object dataset for sim2real transfer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.516242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:22.198067Z digest=sha256:dc31830ab37cf2d235d0757c1dac2414e9176d91ca051f275acd6691a31364e9

Observation 6dacc2bd-93f8-45eb-9a08-28d9e34f710e · outbound

This paper cites See, hear, and feel: Smart sensory fusion for robotic manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction See, hear, and feel: Smart sensory fusion for robotic manipulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.352472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:22.334747Z digest=sha256:976e0c72ff76bca3b738c4dba914a46c2d0700f2602486bd72b464afc27bc828

Observation 2901758c-fe3c-4126-b33d-0b0937768079 · outbound

This paper cites Active clothing material perception using tactile sensing and deep learning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Active clothing material perception using tactile sensing and deep learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.101643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:22.466478Z digest=sha256:01ff23f3adde1472eefa4f53c8102aa188e81469f58a22d9becb9e499e4435ca

Observation 8c16cee2-1899-4b26-ad14-2bd330bd1bca · outbound

This paper cites Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:22.601918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:22.601918Z digest=sha256:c91a2a8db09c11729dd80f6093512850d3dcfcffc66badeb92645f96f864bfbb

Observation beac2993-c198-4208-8264-2b4af7e285c2 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:22.755242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:22.755242Z digest=sha256:a11f98c3dbd48ee1f0fac0125d7e84460596c933efd3e5f84bc4d3d1fa2b0341

Observation eab5930d-261a-4dc9-a0c9-682cbba1406e · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae v2: Scaling video masked autoencoders with dual masking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.809973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:22.850972Z digest=sha256:6ffce17ff9552da8c103a637a3f60d96c616a8ec51ff34b244e14d4c86803726

Observation cdc1cd11-5c75-4fe2-b7cf-614719443e24 · outbound

This paper cites Sigma: Sinkhorn-guided masked video modeling.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Sigma: Sinkhorn-guided masked video modeling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.582233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:22.991207Z digest=sha256:49e9750e71ca082e1d6d61e948b84b66e4b2c52644f5c97ce795c9e9706e9cf4

Observation 2f57f20a-4a5a-4dbb-8f77-630bc069549c · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Mgmae: Motion guided masking for video masked autoencoding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.355424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:23.156250Z digest=sha256:a209e4e3a0c57a6b133a0b0e32f889b6258b9a2bd65753089f804d888e344a8c

Observation 8890e7dc-b217-4dbd-9e61-553736bbc3c4 · outbound

This paper cites VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining.

Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:26.914536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:23.238850Z digest=sha256:5472dd9a78dcbf8d07cc8561f12d6c33276b7aaccbffa6f38743b44d87837924

Observation 9cbbb563-ba2a-4b99-a2b9-4a5c98629509 · outbound

This paper cites Videomac: Video masked autoencoders meet convnets.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomac: Video masked autoencoders meet convnets

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.125395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:23.337796Z digest=sha256:59f208eb42260c8d2e0441f4dd58ae79c745956cb67009c230467e9854ccef84

Observation 6118caa9-3120-4001-9b1f-e5fe9de1fba9 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.432848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.432848Z digest=sha256:bfc81e0ad3aff462d0695a926bc5b4be2bb4b1106d779406b3981747963d7968

Observation 9995e569-a2b3-4ea4-9e92-b43426a67bcd · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Vipergpt: Visual inference via python execution for reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.595338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.595338Z digest=sha256:085f0ac366e0ba088ed7c0aa96428756e23ba1a8be72525065fb494094aeca70

Observation a3013d9a-5e2f-4baa-895f-36b61cf644e3 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.681350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.681350Z digest=sha256:8b982b2e386bf47cb0cda8c230f9a48d090ccef608341f4561b89fa6639814c2

Observation 97709027-603c-493b-a6dc-e93839b22fad · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Lora: Low-rank adaptation of large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.786995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.786995Z digest=sha256:46baa6367229a3b563be558e457fb0f549fca681ed08335da509f1ac13b8557f

Observation 352771ac-e717-4add-8dd5-0cdf78a2e61e · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.868777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:23.865887Z digest=sha256:39e603eca7fd7f1cafcff457321029bf57a4d20b70cc3daefef4e29784e5f33f

Observation 97ba6137-fecf-4e2f-b213-67f56a72c03b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.964185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.964185Z digest=sha256:77edb5f54c3a25d35cb0188eef41b2e4b63a60ee03227b5cfa7ae66d3a65e290

Observation 078e2ba6-a78d-4edd-a9b7-f6b6dd95d72f · outbound

This paper cites Improved baselines with visual instruction tuning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Improved baselines with visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.036960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.036960Z digest=sha256:075c884e05944fef0d307022dee46e9ee78febf96df262128373cc991920a6f2

Observation d76a937d-66db-41a6-9a14-10efc16017b2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.134433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.134433Z digest=sha256:2bab392955551f2f200eec6b3aa0c3292b5f33eb3ab2c0fee9dcc4044b4f2853

Observation 4d471dd7-0154-42c7-9e64-3e2961686ada · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.251376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.251376Z digest=sha256:f970a65b9b9883a4c9b0c004a5ffe86cbb40692252be28aa2f53ddb5ae3fc6a2

Observation d3b73640-6507-408d-a905-1fce1cd35252 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.344367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.344367Z digest=sha256:987d8224de461f56c691a320defee34f42595539bfcb443b69a6e3a0e58f2fe8

Observation 50fa4dfe-7599-43b1-b803-b66e80ffddb1 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.408542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.408542Z digest=sha256:9d115218593f150158bd3f87297186ed36d103ffa58a51a011030987d6e36010

Observation a3f3347f-591d-4657-a834-23fbaa070590 · outbound

This paper cites Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.502357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.502357Z digest=sha256:0b22ba1bc34c1d3e89361cd411844c0cd6bc733b81d36a691f1b94501db36a84

Observation 4426ec64-3622-445e-ae22-47279bff05cd · outbound

This paper cites Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.596314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.596314Z digest=sha256:157daf652089e5106b2b791c5086d0ef03dfd6add83630504be82302a5292a60

Observation 931b65ea-f382-4776-81ba-6b669b203d5d · outbound

This paper cites Cubic spline interpolation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Cubic spline interpolation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.617259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:24.715771Z digest=sha256:846bfc705b64462d04b81581c473fff48ee5cc60db1fd2ce405ab62e9c673cc0

Observation d3d1544c-7f2d-4937-94c2-cc7062024bc7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Universal Visuo-Tactile Video Understanding for Embodied Interaction An image is worth 16x16 words: Transformers for image recognition at scale

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.313082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:24.840567Z digest=sha256:94b83214dc839bb0a14551f409e4b0d43725b1188c4152feb5aba58418fd9f02

Observation efad683c-0329-41d4-8300-4b83db121f71 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian Error Linear Units (GELUs)

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.983127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.983127Z digest=sha256:913d5a975f32e4a9c91d648873addcc82029d6abdaaad1dcfb6cbef2db950869

Observation 6b0cadcd-cdd2-481d-b2d0-2be2fd96087a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Masked autoencoders are scalable vision learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.062498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.062498Z digest=sha256:ef00208ee9e810a38d950a0fe8499849a2d0d513c7e4a750611e5a7acf9c45a9

Observation d9c45fc7-ed47-49be-9795-9d6c2fc851f6 · outbound

This paper cites Gaussian mixture models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian mixture models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.079862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:25.153871Z digest=sha256:d1c7fb9d62194b77a369599f2a34db6ff566d8cea7484a9ff7866ff8cd22bdc9

Observation e15719c8-7422-4632-91e9-557f79f72825 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Raft: Recurrent all-pairs field transforms for optical flow

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.284624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.284624Z digest=sha256:a5369e18f7a361ef8d1d0e606404b537832691c17f450e25606dc32b1aa01302

Observation e6a60081-4ffb-4412-b2f8-1304ccf66c42 · outbound

This paper cites Forward and backward warping for optical flow-based frame interpolation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Forward and backward warping for optical flow-based frame interpolation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.860591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:25.397975Z digest=sha256:e3f401fae308b5a2d536a6692abf426cac5b9590524a5f6add7c9059a358b021

Observation d6bd2af6-b91d-445a-b71f-16d16cb38fe9 · outbound

This paper cites Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.612526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:25.529812Z digest=sha256:2a119ced939f057e2d3add9125ecabb246381860758852fff10bf71342e745d5

Observation 362bd598-074d-4d98-8f7c-4449b1589cb9 · outbound

This paper cites Cross-entropy loss functions: Theoretical analysis and applications.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Cross-entropy loss functions: Theoretical analysis and applications

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.692625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.692625Z digest=sha256:11ec30724265183f1d9388ebcdd15a1023077711b92d4066374d9436c59792b7

Observation b16cc7da-a388-4856-8a25-5582d5a000fd · outbound

This paper cites GPT-4o System Card.

Universal Visuo-Tactile Video Understanding for Embodied Interaction GPT-4o System Card

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.820834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.820834Z digest=sha256:e73b89f28f75b872e50c5f458be90debcc6efe41e3b716c971996f062ef60567

Observation a304cf9d-0f65-4eef-af4e-93345d8f906c · outbound

This paper cites gemini-2.5-pro-preview-05-06.

Universal Visuo-Tactile Video Understanding for Embodied Interaction gemini-2.5-pro-preview-05-06

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.360930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:09:25.923745Z digest=sha256:3e81728b9002448e9d1e05ddb20f2894031882c8dbc8a9314680d343ab8ca676

Observation f3a13097-92b7-47d9-b443-a76e1e1f8917 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-OneVision: Easy Visual Task Transfer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.070117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.070117Z digest=sha256:6e343f5cab7650f75025bc81223d02d3d1af0f8ed1121d6f3ed4f79a93c0a0aa

Observation 120b31ac-4054-4861-882b-b452c3ddd294 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.216112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.216112Z digest=sha256:ccbe123e121e4eaa60385f054bb220e6ab8dcc6a2e25b458e57e266cea1f9751

Observation c51cfa79-1d7a-46af-bebb-b92b42132eb1 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.367593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.367593Z digest=sha256:12f5c9259f287b7baaed5add2f74ce05dfe8a805fd7199a68f7e927ebd67a31f

Observation ee21c921-e493-4641-8370-9e04b885dc58 · outbound

This paper cites Qwen2.5-VL Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5-VL Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.504997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.504997Z digest=sha256:5ff9ccf8fb5f836500d9ae3b5ae0c9fbdf257e1851378c3a8d1879a672890594

Pith citing papers

Observation abc599b5-bf7f-4fd9-bc19-3353ec6d9c16 · inbound

SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes cites this paper.

SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:47:20.035585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:47:20.035585Z digest=sha256:299f19ed40b18ad95973d79c6998ac1996c9cf865e7103508dc27ad269e1e946

Observation 742e49d2-d20c-4ff6-bfd8-ff1b8067ab7b · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:11.718986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:a5393b45fa457183b5fc7787417f624703c68e5e77f5d872e365dd09ce98edf7

Observation 38e50fb9-092a-48b4-b69d-ba9e2c9a760a · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.117773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:f239036555bba4722d452968f7e2a964bf247cd95cfb25d758593344dc924a5e

Observation 0fdca6f4-85a3-4da9-adb4-c103fec5f06d · inbound

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms cites this paper.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.369400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:b131c91d65e9e4f220d30c4875bfe96dc7ebf7357f40e95c92e2b05d5c7b89ce

Observation 34d78103-deb5-4e72-84f6-8f3bac3128f1 · inbound

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation cites this paper.

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.614055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T05:56:03.597839Z digest=sha256:f81e9599fb344c5141708d4492ae783ca2b21982e4c8fa471d7a53247da64e04

Observation 48f125a2-bfc2-4656-84fb-a1c77d22f31a · inbound

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios cites this paper.

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:20:49.438388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:20:49.438388Z digest=sha256:2130ea8ccd16cecbce9164ee71e371b7a50a8fe2061dccd374b6c1d1c75ca699