Pith. sign in

Paper Citation Record · LEDGER

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

As of 15 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2506.06535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06535 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:59:30.182455Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:18:55.243330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:00:10.735031Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy47
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95d3747e-a435-4af7-819f-b34f2fe1c027 · outbound

This paper cites End-to-end train- able deep neural network for robotic grasp detection and semantic segmentation from rgb.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping End-to-end train- able deep neural network for robotic grasp detection and semantic segmentation from rgb

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.563723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:22.608398Z digest=sha256:67b4d7a4dfa85a392d080ec2ddd02f3c93774e73d674d1c04f84eb385a968725

Observation d5230111-0944-4869-aaa7-51d05ad76cda · outbound

This paper cites Hifi-cs: Towards open vocabulary visual grounding for robotic grasping using vision-language mod- els, 2024.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Hifi-cs: Towards open vocabulary visual grounding for robotic grasping using vision-language mod- els, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:22.722777Z digest=sha256:0371c6a1599ea8c77409d11d29c34a0584e81b4a3c35e84e6de0762737adb351

Observation 1d6d9234-71cd-4c17-b891-6ab7bd47d82f · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.284066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:22.879648Z digest=sha256:e541faa1d31d1707ae2b6e69565167da71dc6ac8c5d1c1988b55bb3f365bc25b

Observation 9b9661ce-8956-4d31-a4b7-792933afa43c · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:23.007621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:23.007621Z digest=sha256:91d946ae6dfd2289ebbbf2594aa734e0cfb559a1ff54785a28a77aa116e5f959

Observation 8804535b-9083-4b43-ae69-5573142dff35 · outbound

This paper cites Trust the PRoc3s: Solving long-horizon robotics problems with LLMs and constraint satisfaction.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Trust the PRoc3s: Solving long-horizon robotics problems with LLMs and constraint satisfaction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.116592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:23.129115Z digest=sha256:e4c15cf01b2bd5a85c14f7933580de7545a53815f749c031b2f67741123fc02e

Observation c9c66032-d282-4326-80ec-a5260fcb86e0 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.002526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:23.245485Z digest=sha256:1ff4ec89b125cc1dd743d8ba40a0da65c21fda77c31c32a2c1c1c9dd92405118

Observation 72636aec-7308-464e-8e8b-98f7078e2290 · outbound

This paper cites Jacquard: A large scale dataset for robotic grasp detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Jacquard: A large scale dataset for robotic grasp detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.873290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:23.383235Z digest=sha256:8a3375ffbc4ae45c0e507d004e3c1fa7a3e6a1a3e3372d1f0608217860bbaa25

Observation 4449ba98-7be0-468f-8750-45f386955cb7 · outbound

This paper cites GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:23.566111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:23.566111Z digest=sha256:8b443d8157db515150c785e86b7179e96225be671d909cd890b538fd07bda045

Observation c4544b63-71b8-4fb6-a098-b045da646004 · outbound

This paper cites Acronym: A large-scale grasp dataset based on simulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Acronym: A large-scale grasp dataset based on simulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.729654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:23.734957Z digest=sha256:a40cdeea2d32ddc12fdf8219bf025e213e8ec046fec0b223c2ef8f8fe9d0e775

Observation 0058bcd9-a88b-4582-98c9-09c158d40814 · outbound

This paper cites GraspNet-1Billion: A large-scale benchmark for general object grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspNet-1Billion: A large-scale benchmark for general object grasping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.592702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:23.863902Z digest=sha256:a1e68708b3f5a543eea4f33b150fc99d831ff02cd49901eae5404a1261882e14

Observation 80e38701-26c3-4df1-8cc3-b79c95fa4ef2 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Physically grounded vision-language models for robotic manipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.421739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:23.974250Z digest=sha256:60ef33691cad8b89ab9f22b5d5f2490425dda33ed22a16ad70c729e5a2c386bf

Observation 1deaa556-6bb4-422e-8f02-e061f21cff2c · outbound

This paper cites Rvt2: Learning precise manipulation from few demonstrations.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Rvt2: Learning precise manipulation from few demonstrations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.245120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.079791Z digest=sha256:1e3255e2dbd7ae45ee84b81371b1136ea806fbd244b991f1f89ba65be47bd946

Observation 4377422f-956a-490b-9c78-932e5a1f2b3c · outbound

This paper cites Language-grounded dy- namic scene graphs for interactive object search with mobile manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-grounded dy- namic scene graphs for interactive object search with mobile manipulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.075236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.256736Z digest=sha256:59934032fdb0c1fa4b703dc686ea63ec52820b891b4ef4c6b8da430cfb0cd8f1

Observation 98abafa1-c906-46d5-95e6-45a8c9ce0719 · outbound

This paper cites Inner monologue: Embodied reason- ing through planning with language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Inner monologue: Embodied reason- ing through planning with language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.886801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.387555Z digest=sha256:4ec27479a393d6b6a8b00c7f0464858dd3fc32952966b6ba37df39a99cae88d2

Observation e3a2a822-747f-408d-88b2-8dec101af597 · outbound

This paper cites Effi- cient grasping from rgbd images: Learning using a new rect- angle representation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Effi- cient grasping from rgbd images: Learning using a new rect- angle representation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.701046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.502989Z digest=sha256:869aa8a78b3c623920ff8382d0cffc4e8969b7d63ae63e7f40c06def31be3000

Observation ac6541b7-d66c-424f-9970-9ab555c89a33 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.504139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.611566Z digest=sha256:e10ae4599ffc5b7f38dfbe50eb91ee09867dd69621a002b1cba3291890af7e98

Observation efc502b1-c795-4751-9c80-ca80a1a1160b · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping OpenVLA: An Open-Source Vision-Language-Action Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.715654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.715654Z digest=sha256:cc1ea04237bd3b3da53be54b39e64a35ca3e10bdd1e10d2063d39f54e116f580

Observation 64924a4c-7d2e-4595-8831-24af304a44b5 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.756115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.756115Z digest=sha256:c3e329023a03e720b3d289744bee11d71d5518ca27784369ae1f15b8225aaa22

Observation 8acab195-1467-4563-a6c8-7b772ae19fca · outbound

This paper cites Segment anything.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Segment anything

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.320177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.859095Z digest=sha256:427f604716c85badaff77f85ec6c575747dd0b4eb6312df69e5b834ffa18108c

Observation 6a96d6ae-0ffe-4182-b7f6-2a1a65e858ec · outbound

This paper cites Antipodal robotic grasping using generative residual convolutional neu- ral network.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Antipodal robotic grasping using generative residual convolutional neu- ral network

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.135015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:24.973178Z digest=sha256:fb926e3ea1c2e6ed3cc90401f693a84f9b5e1802e2f354b8a0b90fa50399a957

Observation 8d15859f-698d-4b35-8eca-405fcc3b0c96 · outbound

This paper cites Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.950064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.044852Z digest=sha256:ac2315b8c1eb39c1774e7783e4cf6f4152554bed5aecd0fe1f781b68655e4195

Observation a8950cf3-5ee1-4482-b786-7007bb0e4524 · outbound

This paper cites Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.713433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.210319Z digest=sha256:ed92161dfb7eb24b5ef5548986303968c54aecb03a6660bf28d9513ddc027ecd

Observation 31909571-58c9-41a6-b582-3ab879eff332 · outbound

This paper cites Vision-language foun- dation models as effective robot imitators.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vision-language foun- dation models as effective robot imitators

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.488906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.336778Z digest=sha256:9b3b87459806a38b8bccb7b7b3bb974ab2a9ad0d9de96a836c89bd442dc145fb

Observation 15acdb14-72b2-4a3b-9598-749606a6de8d · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:35.267215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.480480Z digest=sha256:73f411557c1fec6011a1670359cd9db452a5f415b1ca82d36334b5c28b1329a5

Observation 0c3965e4-eaeb-458b-8cd7-3bf30b0f60a7 · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:25.594969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:25.594969Z digest=sha256:bdf3548b6bc4bdd0ac9287fa18b939dc4119ed26985718602b8f73b2bc3915f4

Observation 4e0eb91f-eaf0-42a7-ac46-1c547ca68b6e · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.045677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.698166Z digest=sha256:968b8d3c4c2fd9f0bdd9860b65bed04276630ea2e021a00b665ffa0b674df4ff

Observation 04bbb5e5-bf10-45f8-b658-e3ea8faf660c · outbound

This paper cites Gao, Xi Vin- cent Wang, and Lihui Wang.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Gao, Xi Vin- cent Wang, and Lihui Wang

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.791071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.822106Z digest=sha256:80e2c397c5bc240a0643817ced4ef34f733f66c0589ec147c49d660e8f19d246

Observation 916c6a14-fd01-4e8d-bafa-bbb4aa0e0f7f · outbound

This paper cites Deepseek-vl: Towards real-world vision- language understanding, 2024.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Deepseek-vl: Towards real-world vision- language understanding, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.642108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:25.948312Z digest=sha256:4e422d6b10c83c14abf198408c4e4c9f75950f29994f8c9b4946a1236cb3af84

Observation 2d5e4f87-565e-4bc3-8ba6-19c355082e09 · outbound

This paper cites Hybrid physical metric for 6-dof grasp pose detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Hybrid physical metric for 6-dof grasp pose detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.476271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:26.088842Z digest=sha256:2bcfba76e9a8bb440c84d1a0493136def3f39c958693378a8f892a30d478d5a1

Observation bde316c3-db5a-4528-8d4e-2d604cd6a3e7 · outbound

This paper cites Vl-grasp: a 6-dof interactive grasp pol- icy for language-oriented objects in cluttered indoor scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vl-grasp: a 6-dof interactive grasp pol- icy for language-oriented objects in cluttered indoor scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.294109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:26.196122Z digest=sha256:22f198d7f6449487ecd1d5e6ce7ec085e2ada35f087e79480b3b9fc78fba3f85

Observation 7f678d07-53ca-4bb0-b1bd-5ad1f2bd9083 · outbound

This paper cites GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.297329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.297329Z digest=sha256:4bf3d11d52930d5fcf5ed9d3f821d0bbb10d30efb1dffcf06b759d9b469e8246

Observation 1ff06100-d67a-4dd8-920d-72a05779c937 · outbound

This paper cites Lightweight language-driven grasp detection using conditional consis- tency model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Lightweight language-driven grasp detection using conditional consis- tency model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.143028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:26.411855Z digest=sha256:7d4292f87a7a28bc05c8ba0cbc7391b9db697d9e5fbb09fe6a59e3e536e232c5

Observation e4d91950-9375-4366-97b8-dc1d6c167b02 · outbound

This paper cites Language-driven 6-dof grasp detection using negative prompt guidance.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-driven 6-dof grasp detection using negative prompt guidance

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.962764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:26.526176Z digest=sha256:f9f2858b43231c3c57d6fa1e502f52ad67333cfc25c9025a0974e09b48fd2d85

Observation 4382b3af-21e9-4b5c-9450-43b3bd7c59be · outbound

This paper cites GraspSAM: When Segment Anything Model Meets Grasp Detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspSAM: When Segment Anything Model Meets Grasp Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.631101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.631101Z digest=sha256:e85ddac5a1fb75fdf929b317ff7350d7faffd4f80a1c56f540524b84c57b0836

Observation 84da3673-2a32-4842-8123-6cac36c08c2b · outbound

This paper cites 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.711288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.711288Z digest=sha256:953306ce0948cb6cb5ba73a451c225ee548ea44f3da0dbc6b9ffdd52025c8f5c

Observation ad675a5d-fbb1-496c-a3a1-c096f658dc72 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.875867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.875867Z digest=sha256:b12dcc64f192000e5b7e8db45ad2b1bc09421355b0220ad68cc3a985683b3e32

Observation 874fc6f0-5cf7-43b7-a473-b25a0c20b72b · outbound

This paper cites SAM 2: Segment anything in images and videos.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping SAM 2: Segment anything in images and videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.836594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:26.961017Z digest=sha256:74135af131ca2e2cfca628e53147ed00a2b0d617146c7fd52df1626d8e17d8d6

Observation ba039e28-f965-4619-a1ee-96657594658f · outbound

This paper cites Sadler, Wei-Lun Chao, and Yu Su.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sadler, Wei-Lun Chao, and Yu Su

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:27.082005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:27.082005Z digest=sha256:c13bc6bb12093170dfaaa546bc0367a889fcb9f3862dc56c8e725a9809d517a9

Observation 90556185-0a7f-4436-8081-5fa6887a1a1d · outbound

This paper cites Grasping in the wild: Learning 6dof closed- loop grasping from low-cost demonstrations.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasping in the wild: Learning 6dof closed- loop grasping from low-cost demonstrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.685490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.203026Z digest=sha256:9771b08adc044ff44d5c86c25d9717d354c586f645806ad9a1d8203a163014e5

Observation aeb9fe25-0073-4a06-bcf6-a018aabd0836 · outbound

This paper cites Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.460345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.307342Z digest=sha256:efcbc655ff0e16c3a146b402e6bb76c7e28971d2a9c81c29489227cc3c1514c6

Observation 1ae8b650-b462-4bc5-ae8c-32693fb1adb5 · outbound

This paper cites Foundationgrasp: Generalizable task-oriented grasping with foundation models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Foundationgrasp: Generalizable task-oriented grasping with foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.269395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.457409Z digest=sha256:9d6b7d9e6dc3274e7ec21faea97ce476fae31260bafff7d44854a81f2c52c7cd

Observation 01eadda2-9a00-4ffc-834c-36c28cbcae26 · outbound

This paper cites Mlp-mixer: An all-mlp ar- chitecture for vision.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Mlp-mixer: An all-mlp ar- chitecture for vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.161225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.584816Z digest=sha256:0228951f241dd1eb458357d4cc6415feb1c15ddd2d453a5c5f3d5997f43c00df

Observation ffc1c13a-1cc4-4606-bcef-97a24d047f52 · outbound

This paper cites Towards open- world grasping with large vision-language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Towards open- world grasping with large vision-language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.918336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.686069Z digest=sha256:239efca013a2cacb9855145e452eec93d940644fd055196357443cb7f916e9b8

Observation 6be11674-9bb7-44c0-8bee-7a94c658a422 · outbound

This paper cites Language-guided robot grasping: Clip-based referring grasp synthesis in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-guided robot grasping: Clip-based referring grasp synthesis in clutter

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.698032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.827578Z digest=sha256:e8eb2eec01d14acf08262ec7f8261335a0a36e0a7d2df9afb81387653915a0d7

Observation 43d95ba8-604b-4d8d-80a3-acabd87a550a · outbound

This paper cites Language-driven grasp de- tection with mask-guided attention.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-driven grasp de- tection with mask-guided attention

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.596186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:27.949297Z digest=sha256:acf8742bec1d29ea8c91c08aac706baeb3ec1bd9e18feaecab3981df5230557e

Observation 6e81fb7e-166a-43db-88ca-6573b096d7f6 · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:32.492609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.038044Z digest=sha256:58348f4d94de78af4b38af1e17709cd94b4d194a955adc7cbe5388e6e7bb7faa

Observation 641a213e-9eda-4e76-a227-60a72c8335e4 · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:32.368506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.161716Z digest=sha256:aada40c92747196435a324fadcba95823d7d1a35dba6dfb8f796cb93a0877907

Observation 1af2a83d-7d1c-480e-95d7-9343529fd530 · outbound

This paper cites Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.271887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.290693Z digest=sha256:ae00531006ff035a26ec4c51b906cb935560abac046b46d267485d9b5bc03094

Observation 1a56d105-9007-4492-b106-b077948e6287 · outbound

This paper cites Grasp as you say: Language-guided dexterous grasp genera- tion.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasp as you say: Language-guided dexterous grasp genera- tion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.171500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.362358Z digest=sha256:c24b72271c2a0af3bd514bc2867714a4095d1fc73807f6facd7af541989dd0ac

Observation 294416c8-393b-4a14-829c-747a29c4fe17 · outbound

This paper cites Tidybot: Personal- ized robot assistance with large language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Tidybot: Personal- ized robot assistance with large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.105335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.524037Z digest=sha256:01c030a878392effc755e7c63bc75c2a2eb56e256b6b2139b9cc40ad08102e88

Observation 3ef0dfe0-0ff4-4228-8fc4-dc9fe8b725c4 · outbound

This paper cites A joint modeling of vision-language-action for target-oriented grasping in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping A joint modeling of vision-language-action for target-oriented grasping in clutter

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.069134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.668792Z digest=sha256:4d7d503e5b5dd8d4f5e2d6463c19783c055a5a2c6c0e8831581d6e520b9c5cc4

Observation 6f45701b-fde2-4643-8fcf-3e20fe3915a1 · outbound

This paper cites Instance-wise Grasp Synthesis for Robotic Grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Instance-wise Grasp Synthesis for Robotic Grasping

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:59:30.485178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.814172Z digest=sha256:70a282a07058fe17e9b32388a5b631c1be1cf330d4b20b99ac0e78b8955fdf10

Observation e1fa16b9-38dc-463b-b2b2-f403246a2b57 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Universal instance perception as object discovery and retrieval

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.030130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:28.954588Z digest=sha256:86ef57c65f8619daac1625259cacb4d93d4ad00170c451133f10641b272a6f65

Observation 9f23039d-fbc9-44a0-934d-a1738bf971a6 · outbound

This paper cites Ground4act: Leveraging visual-language model for collaborative pushing and grasping in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ground4act: Leveraging visual-language model for collaborative pushing and grasping in clutter

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.998882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.021121Z digest=sha256:1f681a29b4de48c021ab6b95eacf294e7b2dd78bb02616f0f8fd08dad16da497

Observation 2470b129-8cad-4f0c-a086-f2b7fc481499 · outbound

This paper cites A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.189968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.189968Z digest=sha256:928bbdaa644336d699472af171b9d0ebdc1068fcb8bed4867c0055d4b5cb77e0

Observation ae0cf24b-1a09-40af-8434-edc2065200ef · outbound

This paper cites Se-resunet: A novel robotic grasp detection method.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Se-resunet: A novel robotic grasp detection method

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.973274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.307333Z digest=sha256:de3fb76b7489bf8c0b4fb9921aabae3047e6590954b7d17c8d968efbab319244

Observation 7396091c-f77e-448b-adbd-dbc59406fb45 · outbound

This paper cites GLiNER: Generalist model for named entity recognition using bidirectional transformer.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GLiNER: Generalist model for named entity recognition using bidirectional transformer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.887794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.380199Z digest=sha256:23abece31ece80b6bdfa4721a1c47664976595f2e0e4a74e4a0a95105ca22294

Observation 0434e761-0c96-4fda-9833-7318bb7ef5d8 · outbound

This paper cites Sigmoid loss for language image pre-training.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sigmoid loss for language image pre-training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.742129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.522314Z digest=sha256:75660282b5aa2648120a9e4b15f4686b10622ee70d4f225e0092679521bcd05f

Observation 98014d88-fc38-407f-8456-6851306eaac8 · outbound

This paper cites Roi-based robotic grasp detection for object overlapping scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Roi-based robotic grasp detection for object overlapping scenes

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.517765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.630514Z digest=sha256:ddd070e8f8afabe16814e821949725f8948ee8884fc10cfec59815e0df591270

Observation 6c734f9e-ea84-4ac8-a2e5-80d9356fc318 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.719733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.719733Z digest=sha256:7cf9a59ea88bd51660b244bdec1428726d57c546d6a1c6ad6df5759e519c8393

Observation 37f2c474-a2c0-495b-ac14-5ce4b1aac357 · outbound

This paper cites Language-guided cat- egory push–grasp synergy learning in clutter by efficiently perceiving object manipulation space.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-guided cat- egory push–grasp synergy learning in clutter by efficiently perceiving object manipulation space

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.296924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.848669Z digest=sha256:c2ccf23b155ea36a34e735aa0d3586c62a5ca0abb73746e9eba791f9d07deeb8

Observation ea22c6f7-ecd9-47cf-af37-89e8b7bf70da · outbound

This paper cites Vlmpc: Vision-language model pre- dictive control for robotic manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vlmpc: Vision-language model pre- dictive control for robotic manipulation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.051808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:29.952184Z digest=sha256:6fa2348027d2678fc512e4195db709501194302a2a13f4157265096728e1bbd5

Observation e4fa7e13-3078-47fa-8767-b217232b7b8e · outbound

This paper cites Grasping detection network with uncertainty estima- tion for confidence-driven semi-supervised domain adapta- tion.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasping detection network with uncertainty estima- tion for confidence-driven semi-supervised domain adapta- tion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:30.848781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:59:30.182455Z digest=sha256:6dec96194a6623dc31bfda2bdfd93acb8b0a19e29de1ac78366ef73ae32c005a

Pith citing papers

Observation 953c6cb1-efc4-4e48-9d4e-ea5a098e3913 · inbound

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models cites this paper.

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:00:10.738442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T19:58:19.309634Z digest=sha256:d1a3df95f6a5285616d8d9bac0144732acd00b1f787457df427ec8be8ebcd97d

Observation 405b64ad-7cef-4384-bd3d-931f7f1737c9 · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.243330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.243330Z digest=sha256:3ad43cdfbff3f3711cb52ebf6d42dbc8cda279d992d2dc25aff603687cbe25e0