Pith. sign in

Paper Citation Record · LEDGER

LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2312.17240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17240 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.453526Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.295057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d00bebc9-12d3-4492-97cd-ffaa4de39556 · inbound

InsightEdit: Towards Better Instruction Following for Image Editing cites this paper.

InsightEdit: Towards Better Instruction Following for Image Editing LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:19:22.549062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:19:22.549062Z digest=sha256:9e95524be99c0526f5f7cfa4faeee1b64c97564fb4e3e13ab2132afb952f2d7c

Observation 7296de46-a270-45fa-bfda-aad08c094336 · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.058288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.058288Z digest=sha256:36f75dd8ae4f1971f47b6a0d5e2b9bd93a8d35b2782de4e985d27f6b7a612153

Observation 1a71a723-d1f5-46f7-8e1c-f8af06a7f56f · inbound

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction cites this paper.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.455904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.455904Z digest=sha256:b2c1f007a5aa6e006205efc1a821fb5e64baae4d5a1bd234e5b71a2bd5bdb4b0

Observation 187437e8-6175-4583-8aab-06a929c0177a · inbound

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement cites this paper.

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:31:43.561548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T12:31:43.494099Z digest=sha256:e4959f88e9c4f3710d5fc207ecb6c8a258f12c97763b4735faa56c1e9a7df2dd

Observation 69a0ce21-5fa4-457b-bef2-8b27471f6639 · inbound

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding cites this paper.

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.453526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.453526Z digest=sha256:9ea0681c17c3f6898fbcf66963ab275ae983f0ea9de8c0af63a97cafd399420b

Observation b2e054a0-798d-44af-9595-cd924842f444 · inbound

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation cites this paper.

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:54.267678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:50:54.267678Z digest=sha256:2618056a1ed33fa70c1727341289f052a0fb472be80885b58dec04f4f5e48b10

Observation c5744de0-5db0-4b55-8d63-9d2eaea7ad2f · inbound

TrackVLA: Embodied Visual Tracking in the Wild cites this paper.

TrackVLA: Embodied Visual Tracking in the Wild LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:41.794743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:41.794743Z digest=sha256:ad5ae7b00fa467f5360eb1aa576e53a909b5c2685e9f6ae54028408df8165ae0

Observation 223c34e4-a933-4b81-a9e9-357fc0e6451b · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.966934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:44.966934Z digest=sha256:1df982c0cdf18ca55acafcc062f0e886144a78af0464e6d438c6e5f5f4d1b5fb

Observation 3ec1dc86-b2f3-4920-a928-7395e1de76c8 · inbound

R2SM: Referring and Reasoning for Selective Masks cites this paper.

R2SM: Referring and Reasoning for Selective Masks LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:28.887917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:28.887917Z digest=sha256:b44c9a543c1820b0472a103de10ab9590fc72c70e3e028fea733189a0cbbe585

Observation 2d7ca62d-95d5-4b56-b589-3428429f9c51 · inbound

Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval cites this paper.

Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:17.601173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:17.601173Z digest=sha256:b1ee284ef853b71bb2d787b154ad17716b4e58b1034bdd9fc2de05b0f78bcd71

Observation 222944e9-9eef-4452-891d-bcb190a9e218 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.670233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.670233Z digest=sha256:20659be023e14d52646bc90792b4f7beb047fee2af8176492b505304d6e7a72c

Observation ae5f731d-8a5b-4e55-8359-535d567b62d7 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.959690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.959690Z digest=sha256:51a9d1d3c4ffa97e849a442930209b90cb9b263a097e61d57c9ba60d06de6653

Observation 9ec78942-fa8f-4d43-9ca0-bbefc02526b5 · inbound

Mitigating Object Hallucinations via Sentence-Level Early Intervention cites this paper.

Mitigating Object Hallucinations via Sentence-Level Early Intervention LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:35:32.505195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T08:31:24.173135Z digest=sha256:1d596f8a8ef4f91c5e985d6f7b62fc46c40b42f319db886cf822af693f1d4d38

Observation 61131abb-d351-468e-868e-806cff1112c9 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:55.305583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:55.305583Z digest=sha256:2be83fff841c04efab6fe457007ae2a7bfd1d9dae43919bc5d8f97924f0f7254

Observation 6526a65c-7687-4b8c-a824-684863e2301c · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:59.844125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:59.844125Z digest=sha256:121a16790d5aa4b9832bd027451dae38e97ba7e14bf3e74646da6fefd5d00ef6

Observation 827c5a9c-da2c-4dd7-80b6-f910464b5d2e · inbound

Object-centric Video Question Answering with Visual Grounding and Referring cites this paper.

Object-centric Video Question Answering with Visual Grounding and Referring LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.352762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.352762Z digest=sha256:72e8e3dc8c6da45ceb650a0ad7fc7c457068932ae7d447081f92d43f4c03d23d

Observation 62f8d31d-e6ab-4856-8fba-4b6b33b1b816 · inbound

SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation cites this paper.

SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:37:43.652138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:37:43.652138Z digest=sha256:57bbfffdf47a73680b544d79ad7d2ae71bed67d3b6f483c698248479fad2be5a

Observation b8e6c691-003d-4ddb-9931-4eabec1e20c5 · inbound

MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images cites this paper.

MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T22:08:38.407309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:08:38.407309Z digest=sha256:f0e16d83721df39e5d1c891c39115a8979d76c0f78fcc9dfdb59c642d441cbb7

Observation a4c3445f-89dd-4afc-95ac-78bf5213f0a9 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.918494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:48df534daf7085236aeec021811d0cc307e1b7d213f3b6c83924a93c39fdbe1d

Observation 42dec495-3e45-4b10-b056-02b1f948c5e3 · inbound

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation cites this paper.

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:33:10.089758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T17:31:33.903063Z digest=sha256:3fdfcd71418da950508f5f6654c8f8d3349c3dd688b26930743e44fd46a6e436

Observation 09cc0b98-6fe9-4af3-849a-c95f70d962c9 · inbound

StAR: Segment Anything Reasoner cites this paper.

StAR: Segment Anything Reasoner LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:38.024845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:38.024845Z digest=sha256:24420ae318592b3f4870d57e509bcdc376cbdb231e2f9daad75fd438a8d6299e

Observation a7783ed9-6660-4ed5-a53b-1abea29bc48d · inbound

Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery cites this paper.

Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:49:54.432320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T09:49:10.948061Z digest=sha256:2541d32172aaaa90ff4b664fe8c623b8da71555571fdb1d96f9c55f5d4034157

Observation 4da97aa2-503b-4371-bb13-ba41ad31e781 · inbound

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation cites this paper.

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:21:01.049928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:11:13.376684Z digest=sha256:886b64f02eaa50ba816e02c9a95b917b9c394f9c409553735a4bf9f2c9f25792

Observation a5064ee8-1523-4465-bc78-854d3e1168a2 · inbound

WildDet3D: Scaling Promptable 3D Detection in the Wild cites this paper.

WildDet3D: Scaling Promptable 3D Detection in the Wild LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:26:00.559406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:38:13.336003Z digest=sha256:efc24f92b66ed1607ad275b5e4e09c00257993ebaa2cd42a8ebf9cd3b34f53d8

Observation 0fb09e23-dbbf-4626-9f6d-71729ae88fd2 · inbound

Online Reasoning Video Object Segmentation cites this paper.

Online Reasoning Video Object Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:26:03.169679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:29:03.843441Z digest=sha256:742e620371b292baddcc39bbcda1d4deefaf3978fd3dd32b8092017b94181e65

Observation 4ae23ad6-39e9-4c56-82c5-b0cc1aa60114 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.155620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:7a419d5c093313025edfc712f4f6f8ec9fe408917e4fb368db4ce01f0eba412c

Observation ff21ad11-394f-41f6-a46c-ea23bfa7bf56 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.902806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.902806Z digest=sha256:70f49df64f38e0c8279bb30d617c2607a6e1561a8e77ab256528cb545f9e2a2c

Observation bb6eae89-5126-4dbd-85e7-df2787c76d0d · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.823338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:22114411b73c0ea6eb4f219d9e62dd38f6adc5d987642ed2f015efc43f1e4455

Observation ecc43556-e6ee-4412-8332-5cae09ad7d80 · inbound

WOW-Seg: A Word-free Open World Segmentation Model cites this paper.

WOW-Seg: A Word-free Open World Segmentation Model LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:27:47.990120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T21:23:14.311122Z digest=sha256:a6b847315dc4bdaf0fd9474881f087db2a7184040e2040d4794091fda6e8d3f3

Observation 5bbb6738-9c8a-4fe4-bfba-e452482681fd · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.398084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:42b38f596e208ad465c6dd14a9be9c445f658ba20f3097330d0ee2a51f248466

Observation 5ffc6271-2490-4f0f-8c69-6b8fbb0dbb25 · inbound

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation cites this paper.

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:30:20.640538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:27:15.911342Z digest=sha256:39aaa1e6b8312060ae53f31133de3e0fc3fc51641d8db9effd1e5834cb2c6e72

Observation 07be46d9-400d-43f4-9488-337d668c1e1b · inbound

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation cites this paper.

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.970627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T16:00:40.812912Z digest=sha256:e8ffe10d7a14f432fbaf90e0a29f042149510df3523d137e304f3368aee11f2a

Observation 4ff49ed0-1aca-4349-8b64-8134dcd75b98 · inbound

InstructSAM: Segment Any Instance with Any Instructions cites this paper.

InstructSAM: Segment Any Instance with Any Instructions LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:02.057177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:34:00.442420Z digest=sha256:4fd71053eb277026e90f03048e36fabe36d52baaefcb5aba341267c21867413f

Observation 24910f6d-07cf-422e-b5f3-ece4397050b8 · inbound

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation cites this paper.

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.171607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:40:16.265708Z digest=sha256:7f332403968d2ea37c87bd62563fc171f9ed1b4f8ef5f4a3f0174ad8aa2cc305

Observation 19b7584a-bc32-4046-8aba-13d03e401982 · inbound

MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models cites this paper.

MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.468492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:29:27.353291Z digest=sha256:85c5493061170e5dbf8bdf492be0a2cd30fad36242c8b5e1b37ed1e80c924068

Observation 1d4fe45b-f45e-43ee-b123-c5cadbbc5314 · inbound

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning cites this paper.

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.643650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T17:01:13.745646Z digest=sha256:e0324c4f213a92a1ba89cd88245c80e0fd650882fd36dc7a296d84b6f096ed15

Observation f3eaae0c-7e12-4fe9-b534-50f7d44475b9 · inbound

CABLE: Cloud-Assisted Bandwidth-efficient LMM-based Encoding for V2X Systems cites this paper.

CABLE: Cloud-Assisted Bandwidth-efficient LMM-based Encoding for V2X Systems LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.851192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T21:35:12.989339Z digest=sha256:2f6283b6222795ee9fa0ca7ce3d7f86d7ec9c45caf182da061dd011fa4fd46d9

Observation 0d660b49-7be2-4f8c-aab7-0ee0e76fd5de · inbound

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM cites this paper.

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.303947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:27:45.796650Z digest=sha256:ba36eadfd9e53b9d0444a309ee850398e24087f3b1ea26df318db881193c8a24

Observation 2d52c394-2ed8-4cbf-972e-56ff8d57e683 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.038018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:efed8358b6944aa818aabb960dc18315cd2b30d120660f66db72b15c16728781

Observation c5298ee2-f63a-4f7a-aa66-fbad1f611bc4 · inbound

CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation cites this paper.

CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:36.252137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T09:58:18.387486Z digest=sha256:5e0f969a606f53b2e533afd7425060c6975ceeea157b729dcb95c830bc27b1e4

Observation 578b0448-f727-4d2e-aeb4-fc2f4c601403 · inbound

InstanceControl: Controllable Complex Image Generation without Instance Labeling cites this paper.

InstanceControl: Controllable Complex Image Generation without Instance Labeling LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.101408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:37:41.030752Z digest=sha256:d9abf425d8ca4bbc05ff335013e49cc884a68965dde63ede188622fa8d7c7633

Observation 9fe1f4f0-f427-4707-9a59-8a177d12232d · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 204

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.296286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:6bffb988ea2c9795a480adeb813f57f2f90a0ceec158c3df319f5645f2d40605

Observation ac68ccb2-f914-4178-9ab5-9416b617993c · inbound

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning cites this paper.

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T23:38:50.776533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:38:50.776533Z digest=sha256:e26a52a1f4600e1f6f3cafc9e454c8f292d186760dc5f7239d37a38dd25dcc27

Observation b69baf3e-d8ac-470a-b82b-1b371b2783bc · inbound

ReferTrack: Referring Then Tracking for Embodied Visual Tracking cites this paper.

ReferTrack: Referring Then Tracking for Embodied Visual Tracking LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:00:26.207266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:00:26.207266Z digest=sha256:f556928f2f0271e842ef4bfdebc5af5bf09769706073b7e18276294571f9ae15

Observation 18dc40db-1ea0-41e8-8a1c-5957d02ea833 · inbound

Vision-Language Grounding as Bidirectional Concept Correspondence cites this paper.

Vision-Language Grounding as Bidirectional Concept Correspondence LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.835625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.835625Z digest=sha256:77105b54fb07fe654e9fcc638fafc601836e55b33bf15bc5a9d65959ee0f842b

Observation a210ecc5-89bc-4640-b532-515336f80c78 · inbound

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection cites this paper.

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:54.163754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:54.163754Z digest=sha256:5b4b185c8207ac0ddb6c89fcd751d5f05aa699dc9b182de4306adc11330e520a

Observation 692ac1df-6975-4af9-9015-58f947838de9 · inbound

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation cites this paper.

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:19:11.179065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:19:11.179065Z digest=sha256:946c43b238c3c6946476737c4a20fa53de6f535fb42741a545f05d6af1628132

Observation 547f481d-5ec1-448e-9bba-973dbfa99ea4 · inbound

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding cites this paper.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 291

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:19.125517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:19.125517Z digest=sha256:5236f6205b4d4eb74ae8fcd814cb256ce802914d964b729de1089bd541bd7006