Pith. sign in

Paper Citation Record · LEDGER

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2506.06535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06535 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:59:30.182455Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:18:55.243330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:00:10.735031Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy47
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95d3747e-a435-4af7-819f-b34f2fe1c027 · outbound

This paper cites End-to-end train- able deep neural network for robotic grasp detection and semantic segmentation from rgb.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping End-to-end train- able deep neural network for robotic grasp detection and semantic segmentation from rgb

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.563723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:22.608398Z digest=sha256:686afd46cdaad93dc004582c246088067e180b00c34446ed75aae27c09ed4a9a

Observation d5230111-0944-4869-aaa7-51d05ad76cda · outbound

This paper cites Hifi-cs: Towards open vocabulary visual grounding for robotic grasping using vision-language mod- els, 2024.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Hifi-cs: Towards open vocabulary visual grounding for robotic grasping using vision-language mod- els, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:22.722777Z digest=sha256:8c05b51f0a6c30e6dfdaa86a83f6dac693e7b65adc7ee105fba7b22831db2d93

Observation 1d6d9234-71cd-4c17-b891-6ab7bd47d82f · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.284066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:22.879648Z digest=sha256:ea9fec6dc3dfb7a5a71be4780eb1940db3e03c034236a89266e5a4d6d20ce05c

Observation 9b9661ce-8956-4d31-a4b7-792933afa43c · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:23.007621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:23.007621Z digest=sha256:98820114a928cdf62000fe7790ddab47b726b910903896bad093aeeb021b3640

Observation 8804535b-9083-4b43-ae69-5573142dff35 · outbound

This paper cites Trust the PRoc3s: Solving long-horizon robotics problems with LLMs and constraint satisfaction.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Trust the PRoc3s: Solving long-horizon robotics problems with LLMs and constraint satisfaction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.116592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:23.129115Z digest=sha256:d2eba082df8c593c5db22048df2d923e633557af2f41a4a12c0e9f3889e3e948

Observation c9c66032-d282-4326-80ec-a5260fcb86e0 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.002526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:23.245485Z digest=sha256:4febffd62b2498f6d7aacfdd04fec90c14d2122eebf8293825c0cea6c058fb0d

Observation 72636aec-7308-464e-8e8b-98f7078e2290 · outbound

This paper cites Jacquard: A large scale dataset for robotic grasp detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Jacquard: A large scale dataset for robotic grasp detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.873290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:23.383235Z digest=sha256:96dc1dd440336b0f3f7077bcecfe47b782cb9b3724a906a596b9079f09fca275

Observation 4449ba98-7be0-468f-8750-45f386955cb7 · outbound

This paper cites GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:23.566111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:23.566111Z digest=sha256:709a1e74a51238e7cb5e9d20581f3e5ab6f7d5996594ef4503db8a7895756415

Observation c4544b63-71b8-4fb6-a098-b045da646004 · outbound

This paper cites Acronym: A large-scale grasp dataset based on simulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Acronym: A large-scale grasp dataset based on simulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.729654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:23.734957Z digest=sha256:86a77a45c50dbb1215ce9e7113263d9266b7e4f45e79feb937c6c4d323695808

Observation 0058bcd9-a88b-4582-98c9-09c158d40814 · outbound

This paper cites GraspNet-1Billion: A large-scale benchmark for general object grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspNet-1Billion: A large-scale benchmark for general object grasping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.592702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:23.863902Z digest=sha256:81175722bdf3d3163fa1994cd262111492e52a8329977638df1c83ec4527a68a

Observation 80e38701-26c3-4df1-8cc3-b79c95fa4ef2 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Physically grounded vision-language models for robotic manipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.421739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:23.974250Z digest=sha256:5a62040d0a5319163896e716d2925fcdc7d6363a178184eefa615e8cc5806cb3

Observation 1deaa556-6bb4-422e-8f02-e061f21cff2c · outbound

This paper cites Rvt2: Learning precise manipulation from few demonstrations.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Rvt2: Learning precise manipulation from few demonstrations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.245120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.079791Z digest=sha256:a63e1d053b3aaf9b2895d9f66762d73c8f607a0abff8723cf1fedc8bfb05e2e9

Observation 4377422f-956a-490b-9c78-932e5a1f2b3c · outbound

This paper cites Language-grounded dy- namic scene graphs for interactive object search with mobile manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-grounded dy- namic scene graphs for interactive object search with mobile manipulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.075236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.256736Z digest=sha256:6f6c65fa688a880c9e0b922303427f7dde14e6f63c1410785c54483cfb323e3b

Observation 98abafa1-c906-46d5-95e6-45a8c9ce0719 · outbound

This paper cites Inner monologue: Embodied reason- ing through planning with language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Inner monologue: Embodied reason- ing through planning with language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.886801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.387555Z digest=sha256:61cc343e0fc4b22ebe2957b4e56c758f59665bb5c0c87f0da4ea21dd16b2d8d4

Observation e3a2a822-747f-408d-88b2-8dec101af597 · outbound

This paper cites Effi- cient grasping from rgbd images: Learning using a new rect- angle representation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Effi- cient grasping from rgbd images: Learning using a new rect- angle representation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.701046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.502989Z digest=sha256:616190c353281c1b1a32cf0626c22823a6916e6ef3aa18bc3cda475624c5aeb8

Observation ac6541b7-d66c-424f-9970-9ab555c89a33 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.504139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.611566Z digest=sha256:f8393702a3fdcf686309ac0d47860bc95b7ac3fe371616010b079fe375a3650e

Observation efc502b1-c795-4751-9c80-ca80a1a1160b · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping OpenVLA: An Open-Source Vision-Language-Action Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.715654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.715654Z digest=sha256:5565abdb4e4a0de303ff47b7a6ebf13b052636d4696b5e0028d3a0027c6622d2

Observation 64924a4c-7d2e-4595-8831-24af304a44b5 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.756115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.756115Z digest=sha256:d8492ef7a194ea08b4e3f664bc3a30a0b9aad0dd180873b71b2aa6d75670e52d

Observation 8acab195-1467-4563-a6c8-7b772ae19fca · outbound

This paper cites Segment anything.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Segment anything

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.320177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.859095Z digest=sha256:d3ea889137f8ea5f3f4f9c7b9ba1faea30679ac4fc66e13a590fc96648c78e72

Observation 6a96d6ae-0ffe-4182-b7f6-2a1a65e858ec · outbound

This paper cites Antipodal robotic grasping using generative residual convolutional neu- ral network.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Antipodal robotic grasping using generative residual convolutional neu- ral network

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.135015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:24.973178Z digest=sha256:e7b5bd76071d3cab087c27da3031b61e8ae131aa579da88aa2d7a495892e3ce1

Observation 8d15859f-698d-4b35-8eca-405fcc3b0c96 · outbound

This paper cites Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.950064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.044852Z digest=sha256:0729b11c591d312077c36da8186f4608652a6eb9a5e88949d7832b81072e6fc5

Observation a8950cf3-5ee1-4482-b786-7007bb0e4524 · outbound

This paper cites Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.713433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.210319Z digest=sha256:a9c1f08e53a5c36d94f8be778c9cd77c1f4402c7b41a5d97cd92481daab008c9

Observation 31909571-58c9-41a6-b582-3ab879eff332 · outbound

This paper cites Vision-language foun- dation models as effective robot imitators.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vision-language foun- dation models as effective robot imitators

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.488906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.336778Z digest=sha256:b4b4ea4274e2901a28425611abbbb8e10254f68ad4e2cd3d1fe85de224d60974

Observation 15acdb14-72b2-4a3b-9598-749606a6de8d · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:35.267215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.480480Z digest=sha256:89d6540f383061b8820f9b0e807b3aca6b400dd4cbb5bd70c7f96c6f33e2c0dd

Observation 0c3965e4-eaeb-458b-8cd7-3bf30b0f60a7 · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:25.594969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:25.594969Z digest=sha256:f6ed812e2e1b22311950509617ba1328d5bbb7d4dcf8c53402adea3d5228b702

Observation 4e0eb91f-eaf0-42a7-ac46-1c547ca68b6e · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.045677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.698166Z digest=sha256:26e48fc52145ece19441c6bfabef64d70fb09d178f6d5c353b32604a876a6c73

Observation 04bbb5e5-bf10-45f8-b658-e3ea8faf660c · outbound

This paper cites Gao, Xi Vin- cent Wang, and Lihui Wang.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Gao, Xi Vin- cent Wang, and Lihui Wang

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.791071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.822106Z digest=sha256:005a824cc779ccf1ca181d14ced6baee6beae9d781ed5f925060717d09523e58

Observation 916c6a14-fd01-4e8d-bafa-bbb4aa0e0f7f · outbound

This paper cites Deepseek-vl: Towards real-world vision- language understanding, 2024.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Deepseek-vl: Towards real-world vision- language understanding, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.642108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:25.948312Z digest=sha256:08d3a6c134104cce6240a21afcc6068984a8f12ab007396099cac2f43db79d15

Observation 2d5e4f87-565e-4bc3-8ba6-19c355082e09 · outbound

This paper cites Hybrid physical metric for 6-dof grasp pose detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Hybrid physical metric for 6-dof grasp pose detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.476271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:26.088842Z digest=sha256:8ab9b6610d6bca22c375e091ef777034f2d5e7e2f0bf4f7ab1e2de159049786d

Observation bde316c3-db5a-4528-8d4e-2d604cd6a3e7 · outbound

This paper cites Vl-grasp: a 6-dof interactive grasp pol- icy for language-oriented objects in cluttered indoor scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vl-grasp: a 6-dof interactive grasp pol- icy for language-oriented objects in cluttered indoor scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.294109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:26.196122Z digest=sha256:4047b094ec1cde76198bace3c7b6bbbc426b42978d5aaf87d31bc1465c6a9563

Observation 7f678d07-53ca-4bb0-b1bd-5ad1f2bd9083 · outbound

This paper cites GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.297329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.297329Z digest=sha256:8485709cf55bdbecab2cdee7be64f3b037ee8b1366ddd49c4d849cb037fa48e1

Observation 1ff06100-d67a-4dd8-920d-72a05779c937 · outbound

This paper cites Lightweight language-driven grasp detection using conditional consis- tency model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Lightweight language-driven grasp detection using conditional consis- tency model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.143028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:26.411855Z digest=sha256:b536aae39a36b0bdf0c0d217ed9a203e13f2e66f903e1ff90b6494869a1141be

Observation e4d91950-9375-4366-97b8-dc1d6c167b02 · outbound

This paper cites Language-driven 6-dof grasp detection using negative prompt guidance.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-driven 6-dof grasp detection using negative prompt guidance

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.962764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:26.526176Z digest=sha256:5403c5ad00ecaaea89656b28bf0a510a0ba507aeed7dc956e71ac608aa53699c

Observation 4382b3af-21e9-4b5c-9450-43b3bd7c59be · outbound

This paper cites GraspSAM: When Segment Anything Model Meets Grasp Detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspSAM: When Segment Anything Model Meets Grasp Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.631101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.631101Z digest=sha256:f93f9f969a169433783b64278a23e8e9c2217f47f4287b10786c242539b490a4

Observation 84da3673-2a32-4842-8123-6cac36c08c2b · outbound

This paper cites 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.711288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.711288Z digest=sha256:3b2410dab66b1bd964f28cc489770540717e03634a013f4ae411afa4b7d5e0fc

Observation ad675a5d-fbb1-496c-a3a1-c096f658dc72 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.875867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.875867Z digest=sha256:cf19b25efba8658fec9e7052ae4cba13cba24d6b13c8812b28d02622fcd01501

Observation 874fc6f0-5cf7-43b7-a473-b25a0c20b72b · outbound

This paper cites SAM 2: Segment anything in images and videos.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping SAM 2: Segment anything in images and videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.836594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:26.961017Z digest=sha256:97c10f1a9fd0562ed51059c2c9e8f6ea8ae4dc3de329e4dfa450b5b468d525a9

Observation ba039e28-f965-4619-a1ee-96657594658f · outbound

This paper cites Sadler, Wei-Lun Chao, and Yu Su.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sadler, Wei-Lun Chao, and Yu Su

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:27.082005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:27.082005Z digest=sha256:4f2ea5a615a012c3792858eb7ed881e974954e474091429e67f5c66cf6a7c9e3

Observation 90556185-0a7f-4436-8081-5fa6887a1a1d · outbound

This paper cites Grasping in the wild: Learning 6dof closed- loop grasping from low-cost demonstrations.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasping in the wild: Learning 6dof closed- loop grasping from low-cost demonstrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.685490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.203026Z digest=sha256:6f286bb1f9bd2fe5d55ebc838aa5fe7ec75fa45201ba0e97a603b558745e0523

Observation aeb9fe25-0073-4a06-bcf6-a018aabd0836 · outbound

This paper cites Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.460345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.307342Z digest=sha256:3fee58ab3ce4a2869cdcf04893247b5d9e5bcf66029f3c10cf462d5e075ad7f1

Observation 1ae8b650-b462-4bc5-ae8c-32693fb1adb5 · outbound

This paper cites Foundationgrasp: Generalizable task-oriented grasping with foundation models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Foundationgrasp: Generalizable task-oriented grasping with foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.269395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.457409Z digest=sha256:c332ce11846ac60a885e4f7c29fb36613e8e03a325128ae7590fabc59cb1bfa8

Observation 01eadda2-9a00-4ffc-834c-36c28cbcae26 · outbound

This paper cites Mlp-mixer: An all-mlp ar- chitecture for vision.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Mlp-mixer: An all-mlp ar- chitecture for vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.161225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.584816Z digest=sha256:ee57cff9f42abfeceffe918ba2628253546adaf658525c2e426e7d7a1fa36860

Observation ffc1c13a-1cc4-4606-bcef-97a24d047f52 · outbound

This paper cites Towards open- world grasping with large vision-language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Towards open- world grasping with large vision-language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.918336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.686069Z digest=sha256:37dc7457d93f732b333e481b9af9ee120708aeebda6316f2f7aa8725af93b5d0

Observation 6be11674-9bb7-44c0-8bee-7a94c658a422 · outbound

This paper cites Language-guided robot grasping: Clip-based referring grasp synthesis in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-guided robot grasping: Clip-based referring grasp synthesis in clutter

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.698032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.827578Z digest=sha256:06a296c75e9874737117af5ab80d2504fae1f21706fc782b85b85fe5eb229bd4

Observation 43d95ba8-604b-4d8d-80a3-acabd87a550a · outbound

This paper cites Language-driven grasp de- tection with mask-guided attention.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-driven grasp de- tection with mask-guided attention

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.596186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:27.949297Z digest=sha256:37c1967c7a8df57b4b9aa5c18bdbd23001cf241e152304dfb759dc5c252991a5

Observation 6e81fb7e-166a-43db-88ca-6573b096d7f6 · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:32.492609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.038044Z digest=sha256:f82fac982f875d95dddf3b3508fac8ba529bb9d30589682301b72f1df908b5f8

Observation 641a213e-9eda-4e76-a227-60a72c8335e4 · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:32.368506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.161716Z digest=sha256:54b5895bb2736cd168112236366192b09e14d6c3c31c3efb99de96a50cf0b719

Observation 1af2a83d-7d1c-480e-95d7-9343529fd530 · outbound

This paper cites Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.271887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.290693Z digest=sha256:ad6a67037401703bdce2b4661cb9f836d2b8a881cfc9fbe0080f23df653907a3

Observation 1a56d105-9007-4492-b106-b077948e6287 · outbound

This paper cites Grasp as you say: Language-guided dexterous grasp genera- tion.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasp as you say: Language-guided dexterous grasp genera- tion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.171500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.362358Z digest=sha256:2bc279bb3a5053b38a68f2d34f0aef920f59dc8bffb6cc884bd4b58a212db029

Observation 294416c8-393b-4a14-829c-747a29c4fe17 · outbound

This paper cites Tidybot: Personal- ized robot assistance with large language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Tidybot: Personal- ized robot assistance with large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.105335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.524037Z digest=sha256:1852d5863018b91e90a3a21e3bf3e530b3d2498a56690af6b77b7a38c234636a

Observation 3ef0dfe0-0ff4-4228-8fc4-dc9fe8b725c4 · outbound

This paper cites A joint modeling of vision-language-action for target-oriented grasping in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping A joint modeling of vision-language-action for target-oriented grasping in clutter

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.069134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.668792Z digest=sha256:3ce81ed24d6279419d58f97531c95665352be729abecb52d64944997a9beeab1

Observation 6f45701b-fde2-4643-8fcf-3e20fe3915a1 · outbound

This paper cites Instance-wise Grasp Synthesis for Robotic Grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Instance-wise Grasp Synthesis for Robotic Grasping

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:59:30.485178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.814172Z digest=sha256:941953e96cfde7533f286c3b96c5eb30a4f6209a60ae4664405042e061acf013

Observation e1fa16b9-38dc-463b-b2b2-f403246a2b57 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Universal instance perception as object discovery and retrieval

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.030130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:28.954588Z digest=sha256:48354c8d8e52d2ffed9ce2c30624d8ef75c5762e4fc89dae6c50246689977679

Observation 9f23039d-fbc9-44a0-934d-a1738bf971a6 · outbound

This paper cites Ground4act: Leveraging visual-language model for collaborative pushing and grasping in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ground4act: Leveraging visual-language model for collaborative pushing and grasping in clutter

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.998882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.021121Z digest=sha256:4dcacaad9ecac9b4e66ca6d013ec87b30b12fc7ac3832197451265e565ca94c6

Observation 2470b129-8cad-4f0c-a086-f2b7fc481499 · outbound

This paper cites A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.189968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.189968Z digest=sha256:56a74ccc4019a11a6f37bc9870b87c9208d74f42186ca4d00741cf08f9642a70

Observation ae0cf24b-1a09-40af-8434-edc2065200ef · outbound

This paper cites Se-resunet: A novel robotic grasp detection method.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Se-resunet: A novel robotic grasp detection method

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.973274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.307333Z digest=sha256:f954978c9d328ca9ad238964363817120b988ca564e2cd406c9ecc1e7a13b999

Observation 7396091c-f77e-448b-adbd-dbc59406fb45 · outbound

This paper cites GLiNER: Generalist model for named entity recognition using bidirectional transformer.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GLiNER: Generalist model for named entity recognition using bidirectional transformer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.887794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.380199Z digest=sha256:f3b234bde2ca440ce5e684896fdf44552a094420d89b946bc307820bf022910f

Observation 0434e761-0c96-4fda-9833-7318bb7ef5d8 · outbound

This paper cites Sigmoid loss for language image pre-training.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sigmoid loss for language image pre-training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.742129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.522314Z digest=sha256:b8471e17422a19a38142b35de9fdd667cf9de188731f508f5b502c9fcce07105

Observation 98014d88-fc38-407f-8456-6851306eaac8 · outbound

This paper cites Roi-based robotic grasp detection for object overlapping scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Roi-based robotic grasp detection for object overlapping scenes

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.517765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.630514Z digest=sha256:5ec8374668f025164d1862c648b4c777ebb04cd81438f334b15ce1e038b4e12e

Observation 6c734f9e-ea84-4ac8-a2e5-80d9356fc318 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.719733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.719733Z digest=sha256:4d5c4827c697acf14d24bb7db83c9814e41f92c92905cb6808d627600dd83eee

Observation 37f2c474-a2c0-495b-ac14-5ce4b1aac357 · outbound

This paper cites Language-guided cat- egory push–grasp synergy learning in clutter by efficiently perceiving object manipulation space.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-guided cat- egory push–grasp synergy learning in clutter by efficiently perceiving object manipulation space

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.296924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.848669Z digest=sha256:7c3c915ae9aee2cc74cc0e634278b65e2eca38d0af6589666e69567f41917347

Observation ea22c6f7-ecd9-47cf-af37-89e8b7bf70da · outbound

This paper cites Vlmpc: Vision-language model pre- dictive control for robotic manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vlmpc: Vision-language model pre- dictive control for robotic manipulation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.051808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:29.952184Z digest=sha256:914cca25072b3cfa0b6591b62ff7a800a68c6e36774ed517d8c2749520a8d1be

Observation e4fa7e13-3078-47fa-8767-b217232b7b8e · outbound

This paper cites Grasping detection network with uncertainty estima- tion for confidence-driven semi-supervised domain adapta- tion.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasping detection network with uncertainty estima- tion for confidence-driven semi-supervised domain adapta- tion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:30.848781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:59:30.182455Z digest=sha256:746932ec400157dd7d1c2cf64dd8a46b1ec4fca83794e755b556249a231a858e

Pith citing papers

Observation 953c6cb1-efc4-4e48-9d4e-ea5a098e3913 · inbound

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models cites this paper.

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:00:10.738442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T19:58:19.309634Z digest=sha256:b56671ecacb1d04c27dff9396f8a3bd57fc6a8efc7493ec3a23da0279fb22042

Observation 405b64ad-7cef-4384-bd3d-931f7f1737c9 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.243330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.243330Z digest=sha256:257c151c973db7ade61bd01521733302abe02257db7ef89bb327d3674937dbea