Pith. sign in

Paper Citation Record · LEDGER

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

As of 14 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 2 inbound Pith citation observations for arXiv:2412.00153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00153 v3

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:10:41.196244Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:43:32.954256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.913171Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc9054c9-867b-423a-b426-2fa2669f5baf · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.602727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.602727Z digest=sha256:4395282f95f9a4a08c6817e541995f63468b91d417f1142a811ef573e4d1d5cf

Observation 33724053-7262-4073-b3cf-99c545c9044c · outbound

This paper cites Barron, Fer- ran Marques, and Jitendra Malik.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Barron, Fer- ran Marques, and Jitendra Malik

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.608496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.608496Z digest=sha256:0d897bcf313cb15dc116559f2f39a4ac0086b7cc9b950ce5f4cdd7097801a374

Observation 406c57ea-401b-4377-82dd-566577064de0 · outbound

This paper cites Qwen Technical Report.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.614159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.614159Z digest=sha256:92cf82308aebaf3eb2863f5169e6bac9c43097ef38db3070664f8e872879bde2

Observation 8e606117-7d9f-443f-bb03-db732569e074 · outbound

This paper cites CoReS: Orchestrating the Dance of Reasoning and Segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model CoReS: Orchestrating the Dance of Reasoning and Segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.619538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.619538Z digest=sha256:be0ffb85071c9f63e89ac089ec341e372ec10067af59b4d177c0425d87c62d36

Observation 8670c5ab-c675-4e1c-a7c2-8de69265d678 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.625056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.625056Z digest=sha256:9a5f1ab164be3d4f481b46021a0a25aa9b2084b5154d6ceeaa18f9504877195e

Observation baea3694-043d-48ac-8e92-fe3d7526c3b3 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Coco- stuff: Thing and stuff classes in context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.630566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.630566Z digest=sha256:d33ac10e182c8574e1aafe44b09aa09a7c587e2b5e6df3614d20bf8b25081eaf

Observation a4dae4b5-f4a4-496e-acb3-bec70a780927 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.636951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.636951Z digest=sha256:40cb80dba8bc714814853d01582c6867aee5f3a65c02075e6e8f0aa7f1559e42

Observation 01ddf6a1-4087-4e6b-96da-84fe45560491 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.646884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.646884Z digest=sha256:13f574417c6a60109ad66156ac29ebd520be7cf1b4262be521877a384e9f9b68

Observation ee8b1da5-7b65-49ca-b49a-e87c59f6ab99 · outbound

This paper cites an unresolved cited work.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.655072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.655072Z digest=sha256:ab44fdb08c47c35ff3237b357dee11409bb711cd951c2ed65d2532d35db076a4

Observation 7c3a0701-eb07-4aa1-a8e6-ac61d243cd8a · outbound

This paper cites Rethinking Atrous Convolution for Semantic Image Segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Rethinking Atrous Convolution for Semantic Image Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.663763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.663763Z digest=sha256:7ca0f434072af913c732271ccd35bb9007b3c6d4f29c79f696d7c21d5f597111

Observation 63341173-fefd-42f3-bdce-6bfaa1798857 · outbound

This paper cites A Unified Sequence Interface for Vision Tasks.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model A Unified Sequence Interface for Vision Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.671530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.671530Z digest=sha256:e66cd1f801b527ffafb750d49dc595c2c5be1bc7cbc049a6f05b3d3bdb0d92ce

Observation b7eb7c55-f6dd-4d02-bb83-2848e14e3f8e · outbound

This paper cites Schwing, and Alexander Kir- illov.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Schwing, and Alexander Kir- illov

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.678258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.678258Z digest=sha256:db18903c6a26070e7245135fb03bbacefd5d3a260dfb561edd5541c9a2b48510

Observation a79df8a7-3099-4d6b-b159-806618062944 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.684396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.684396Z digest=sha256:f1eb6e7e663505825b70368924be4e4c79945adee0f9928c4118bcffc74d1e7a

Observation f8a1f219-05f4-46f2-8f6b-6705ee8b53ce · outbound

This paper cites CascadePSP: Toward class-agnostic and very high- resolution segmentation via global and local refinement.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model CascadePSP: Toward class-agnostic and very high- resolution segmentation via global and local refinement

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.690426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.690426Z digest=sha256:39fed60650ee5307928bab387765219875a4b9a7c047c5080c2501e33fbfb3b2

Observation 472714cf-0a71-48e2-ad16-e20305fe19bd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.694974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.694974Z digest=sha256:c622eec699774e4eb7566985b52747a49d30a3670671b478be3a4e9c9b5335bb

Observation 0cf0912c-d82e-4403-96eb-81220347d1e6 · outbound

This paper cites Instance-aware se- mantic segmentation via multi-task network cascades.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Instance-aware se- mantic segmentation via multi-task network cascades

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.700692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.700692Z digest=sha256:4f6177f57d563c422cfa51077f126407c5b35826031d9e5b5118daad02f9845a

Observation f9fde1d8-d406-4036-b6e1-d06598d11d26 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.708748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.708748Z digest=sha256:e05f58c94221ac1ba8911a6bba62826a548b39a1c25e84140429f2ebbe778f15

Observation 9b38e67c-72ca-45f5-9857-f083d57431a7 · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Vision-language transformer and query generation for refer- ring segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.720146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.720146Z digest=sha256:742ade1e361d7680ffa171ff35bd3779368a03df7ebcc02ba51555280e1b720b

Observation 86d48a15-4367-4cad-b139-7a0fd8be5078 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model De- coupling zero-shot semantic segmentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.725466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.725466Z digest=sha256:fac98ff1681810d1474f6cd08daae22730b9a363274cd9e6093cf2f3729acc5e

Observation ce5e425e-f964-4565-afbd-cfd6c5d71098 · outbound

This paper cites A discriminatively trained, multiscale, deformable part model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model A discriminatively trained, multiscale, deformable part model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.732244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.732244Z digest=sha256:700b01c4dae99d360c9c15f3870b048024bceb846cbafd156a6835ebfdbc0f78

Observation 712ad1f6-051a-4afe-b9ae-0cf631fd18fa · outbound

This paper cites Dual attention network for scene segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Dual attention network for scene segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.738344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.738344Z digest=sha256:c765721e80165e0ebf4dbfa45596f644c3a9a49edebbf3476dfce0ce717a3387

Observation 468bd23d-d2f5-4d5f-a882-d82b4d1655ea · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Scal- ing open-vocabulary image segmentation with image-level labels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.744388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.744388Z digest=sha256:d84a77acc79ab33a5faa3706c3338067822a54e3699c9b9be6a894d57b591c82

Observation c3f0ae87-5a05-4d50-a933-63407e84642a · outbound

This paper cites Fast r-cnn.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Fast r-cnn

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.829966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.750583Z digest=sha256:71a0650cd8a2488a0de1bceeb9c1dcd511e630004d8d3084cac74a115090272b

Observation 4e1b00fc-59f0-437c-8b0b-7adb9f285eb7 · outbound

This paper cites Efficient hierarchical graph-based video segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Efficient hierarchical graph-based video segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.813092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.757926Z digest=sha256:10712b9e55229600f37270582d7b22ce949e157d6b4b1e7773736b060619357f

Observation feadfcc1-fa40-435e-9f00-49e411c620cf · outbound

This paper cites SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.764202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.764202Z digest=sha256:3407b31a6a08b57af8ce2b47eb0c92ed5874ff86ea151f3bb4a26c5144a74a6a

Observation 3d406e68-77a8-4547-8ecd-a6be74544aa0 · outbound

This paper cites Global knowledge calibration for fast open-vocabulary segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Global knowledge calibration for fast open-vocabulary segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.794206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.770557Z digest=sha256:40cb91b52d5263825afc46788e3b468787b28a5671ebfd8d3cc6333d358b2025

Observation 695588bf-8ac2-4c5f-b90a-4ce4999a5c09 · outbound

This paper cites Multi-modal instruction tuned llms with fine-grained visual perception.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Multi-modal instruction tuned llms with fine-grained visual perception

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.772644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.779725Z digest=sha256:ab1f66e03d7727e15fb39dcdf415ea0ff6120c438e2b2cde5462ef6adeef3ec0

Observation 8328d994-4c35-4f3c-855e-87f1083541ca · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.787127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.787127Z digest=sha256:8d10cb8dbf460de77904cf1cc3c59b9418232d40bdfd396513c8d6f52eebad7b

Observation 667a8c61-474c-4eb2-bdc1-a78bf072d809 · outbound

This paper cites Bi-directional relationship inferring network for referring image segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Bi-directional relationship inferring network for referring image segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.751208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.794615Z digest=sha256:b8429abcfa2f114ddb0fb3779fabc0984358e5172684043720600c65bb1152da

Observation f72b2860-a780-43f2-94a4-44a5abbbbb04 · outbound

This paper cites Referring im- age segmentation via cross-modal progressive comprehen- sion.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Referring im- age segmentation via cross-modal progressive comprehen- sion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.727435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.801817Z digest=sha256:802316402bc9c4b65d19fd988ccf269e47eaae71d7be2e52f4ebc2fa3f2dc577

Observation 8592435d-0e8b-42ed-a9fc-ac42b5963117 · outbound

This paper cites CCNet: Criss-cross attention for semantic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model CCNet: Criss-cross attention for semantic segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.699887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.807249Z digest=sha256:b4195294aa5496fccbe6fac29dfdcfa1d549e3ceb4236fc9c6964eb9cc1ec2fd

Observation 9b3796b3-aed2-4e82-aa22-ac393def43ef · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Scaling up visual and vision-language representation learning with noisy text supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.817090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.817090Z digest=sha256:086bfbcf5052b9ccf10c38e7e116b52c4fee93ca8ce6bce6e5b62cea5a01e218

Observation f72a46c9-8ce4-4987-be4d-80882b00fa95 · outbound

This paper cites Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.655968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.825753Z digest=sha256:620a8265fe309ad584cd627ea1372bb9ebbe382574664c89e8241be6cfad5985

Observation 7fdc2374-2927-4d20-89ef-ce062cd36847 · outbound

This paper cites Locate then segment: A strong pipeline for referring image segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Locate then segment: A strong pipeline for referring image segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.634113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.830882Z digest=sha256:c28c412466e452003649a055de0a7617cd4bd07c26c0b31a9432d8944c3e15ae

Observation 06feee68-e2b9-4bf9-a4f7-05bc6e827007 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.613695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.837314Z digest=sha256:3fd5f34856ef16a1ae788330332c99570263992edd673d03f354d11ef3af2a7f

Observation 6433f450-c112-4599-a060-8f5da956b3be · outbound

This paper cites Segment any- thing.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Segment any- thing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.593290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.845266Z digest=sha256:a799cff1b2b920fde6f7b507e5acd9bc5a5f90af3c503adb68a5afb4ee41723e

Observation b2c2b052-bab4-4579-a33b-62edd58b93a9 · outbound

This paper cites Large language models are zero-shot reasoners.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Large language models are zero-shot reasoners

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.569494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.852552Z digest=sha256:3cddaf8e85d1c188cebfa0b138f91392180575178e2d6abfdda8590e855ff1c8

Observation 83edde41-c518-440b-80c9-ed98b8846423 · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Lisa: Reasoning segmenta- tion via large language model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.859899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.859899Z digest=sha256:6e171d5bf34c055b7277c7936a83dfd9b7a8ca73f378f11d1f0a7b4bbabd939e

Observation c36fba00-d6af-4aeb-86a7-9886842fa00a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.867634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.867634Z digest=sha256:e386e0fe635320a6077e9603d02a2949d58ecaf973fb6ae3703641608d98b7eb

Observation 877046cf-543e-44f7-9091-4932d560821a · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.529184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.873795Z digest=sha256:e7a623ebdd2c7f76175ec0f101957557e54469f0b2eed242f4377b997280f50a

Observation 690077fb-17d7-47a2-b76e-8a6bd67a5223 · outbound

This paper cites Referring transformer: A one- step approach to multi-task visual grounding.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Referring transformer: A one- step approach to multi-task visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.509319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.879409Z digest=sha256:752f5d89bc8e9ac895fedfb3c4ed80304431817581ebca528a0bcc08b52ac5f4

Observation 72a2b959-46dd-4dea-a489-e37fced6d16b · outbound

This paper cites Textbooks are all you need ii: phi-1.5 technical report, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Textbooks are all you need ii: phi-1.5 technical report, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.486608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.885283Z digest=sha256:3cfd2d8f3139404200642118ce5c8dbdabf05970d39c7c89f03373b389ea2307

Observation dec07de0-38ab-4eaa-ac37-913a67839e81 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Open-vocabulary semantic segmentation with mask-adapted clip

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.452323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.892528Z digest=sha256:771c16ea26986e0aaf87a45d2acc804a5a12ef82b2ba58a9113494853922e232

Observation 07ea0b25-15aa-4478-8a64-45321df358fc · outbound

This paper cites Microsoft COCO: Common Objects in Context.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Microsoft COCO: Common Objects in Context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.897561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.897561Z digest=sha256:db73ff220449d1139b86ac1334446cd97a268a11dde6a2b16a3f4a9aae9d90e8

Observation 32b269f7-7bff-4b8b-9bff-f80d8a639b46 · outbound

This paper cites Visual instruction tuning, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Visual instruction tuning, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.903182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.903182Z digest=sha256:cc608685558fa736cbf86d251c56fb48ce9b10bf0f01628005f9044d5719f47a

Observation 12d2280a-d8f1-466d-9559-30e2cbbcb6aa · outbound

This paper cites Fully convolutional networks for semantic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Fully convolutional networks for semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.410159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.908369Z digest=sha256:afbae6edd40adc415f2049c0bbf32fe9c18d5ac2fe53be06591ed9bf9b778f8f

Observation 6e189871-22a0-43f1-b894-35761da8fbf6 · outbound

This paper cites Decoupled Weight Decay Regularization.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Decoupled Weight Decay Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.913998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.913998Z digest=sha256:b0bffb783b7f53e3522b347be5a4a1b10cf68382004065b024ddd0ac3efe608e

Observation c2ab0f93-7bd6-42cc-91df-66e21610c6b1 · outbound

This paper cites Cascade grouped attention network for referring expression segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Cascade grouped attention network for referring expression segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.388448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.919613Z digest=sha256:cc869d5d9b5b3625db3058381177847c485c10287e1d94ea109b6a9c5a2e0373

Observation 104f859b-3744-4b76-8055-2c3b3d9115d3 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Generation and comprehension of unambiguous object descriptions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.366413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.924413Z digest=sha256:ebfb3269468912c74128ba4425897bdd220e758eb8fae1127e985e96613600ac

Observation 872b78b8-0097-48b5-9244-de8b596374c7 · outbound

This paper cites The mapillary vistas dataset for semantic understanding of street scenes.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model The mapillary vistas dataset for semantic understanding of street scenes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.930956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.930956Z digest=sha256:9304a084414137f33bf1396ab82a01ca98ccc9bf57c6869db78d5b59e18e30f7

Observation 019f3a6a-ded3-4691-9e91-19bfd3467ebf · outbound

This paper cites Chatgpt: A language model for conversational ai.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Chatgpt: A language model for conversational ai

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.319108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.935743Z digest=sha256:854a6957f6fa350125625598b7fc936b55f5056001fe86bc958e0ba2656889e4

Observation dac3c2a1-6ded-4d4f-8369-3e8dc3e16b12 · outbound

This paper cites Gpt-4 technical report, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Gpt-4 technical report, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.944616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.944616Z digest=sha256:35f52825627e452797644173335c4188627b8b2d4bd388cf47f433fc10d1eba5

Observation 34c97eb1-eb05-4fbc-b68e-a2a15804a35b · outbound

This paper cites Kosmos-2: Ground- ing multimodal large language models to the world.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Kosmos-2: Ground- ing multimodal large language models to the world

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.278398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.951145Z digest=sha256:00800d5b64f385c4d5aa98b11aa448d70fc0496adf79985e298250461aa4629e

Observation 39d0a7dd-ec4e-4fa7-82c8-81ee1552a315 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.957551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.957551Z digest=sha256:a995df4114cdc6ed948426187126a4b3da41c91872f2bb1d52a4ca412368a67d

Observation 4484be1f-82de-4194-8d0e-62c05722f88a · outbound

This paper cites an unresolved cited work.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:10:42.255454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.963755Z digest=sha256:dcd1785a3ca1e7884ef88fe717488baca11e3f65f03c5b9e1170deee04445704

Observation 770cf0eb-9593-48f3-9639-6196fc696c0a · outbound

This paper cites Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Doll´ar.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Doll´ar

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.235786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.968878Z digest=sha256:47bf7588ccb6247b63b7b39b46735e426aab84fcc4e871bdf8dae050ac14ae2d

Observation c36e3f83-f330-46be-a92a-f67190285945 · outbound

This paper cites Multiscale combinatorial grouping for image segmentation and object proposal gener- ation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Multiscale combinatorial grouping for image segmentation and object proposal gener- ation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.217544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.974233Z digest=sha256:6c34af02c325872aa8cbd0df863dbdca5229d7a3774e24ff22cebc35cfa12ff0

Observation 57339680-1a03-4d27-85a4-2380a7780b00 · outbound

This paper cites Learning to segment every referring object point by point.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Learning to segment every referring object point by point

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.197893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:40.980144Z digest=sha256:5746940ed019ec1c6fd7153f7db1bae1ed340f7ef64e76cdfaba9198f7240b84

Observation 59a4f07e-f360-429a-b3d5-51a7e8aa0aeb · outbound

This paper cites Learning transferable visual models from natural language supervision.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Learning transferable visual models from natural language supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.986159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.986159Z digest=sha256:b5ff92d6da6c31708530535e2df4886790d63c1bb838a95d1fe76e443e76e0f2

Observation d5f58f64-088c-4ebc-88cd-82efb08cd80c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Learn- ing transferable visual models from natural language super- vision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.993676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.993676Z digest=sha256:6d6c69cc8626a84f429cb838ad730e51726b700fbc0c434da1be3ffa419e43df

Observation 95f21376-6334-48a6-8ff1-4d766b914a40 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.129623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.001120Z digest=sha256:28eb0e2b5f188ef3376c22d22d2f3e38b57003100cba6847a9c594542da0fb33

Observation 3e3e2a8f-2770-47de-9ed6-c23a7c8d921d · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.104654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.006137Z digest=sha256:c368be460ca76c033c816f9726dfdd96e4454589f8a69b5fffc39c654ce38c38

Observation 6084e481-980e-45df-863d-ead855e42446 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Pixellm: Pixel reasoning with large multimodal model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.078881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.011579Z digest=sha256:1e0a406f03067ab0f936eb480b71820eb9751f3b1141411a929f7fd4ade7cec1

Observation 091d006e-4246-4292-9933-ae0e13200fe1 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model U-net: Convolutional networks for biomedical image segmentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.050276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.017480Z digest=sha256:f3bebea3d3a5093016be4d055d41d3d5f074b09e11b896bb43706137fe2635f7

Observation b01dd00a-2855-4aad-ac96-e22bd1db2263 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model LLaMA: Open and Efficient Foundation Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.022299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.022299Z digest=sha256:a652af7992d6f7796e0c5539366cd38fc68d7e5e7d08eebf42f3aef8718b761d

Observation efaf51df-9525-4fd6-ae6d-369540f4f34d · outbound

This paper cites Selective search for object recognition.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Selective search for object recognition

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.024766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.027703Z digest=sha256:1e4a538fde0c25927608690c3a0198e79582dee97ced356b1f4d9be170cdbc0f

Observation 4c2effc2-46d7-4be0-96d6-1d5b96e2e933 · outbound

This paper cites Llm-seg: Bridging image segmen- tation and large language model reasoning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Llm-seg: Bridging image segmen- tation and large language model reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.002608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.034013Z digest=sha256:46a458eaade14dfa329c678fba387eb042e8367376c60a1ae8a18d6f8040a395

Observation 2a539345-3b75-4d8e-b1c7-e22198fa5d20 · outbound

This paper cites SegRefiner: Towards model- agnostic segmentation refinement with discrete diffusion process.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SegRefiner: Towards model- agnostic segmentation refinement with discrete diffusion process

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.985143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.040649Z digest=sha256:d035fef906018365ee280585d6ad6ef7e53e377b810e261647b819e4ec66da0f

Observation 9096e147-9792-4ea6-8cbb-84a8b06d9cec · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.049912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.049912Z digest=sha256:ab4624f5e752081013569dc80b1392a4c3227e29bb5a0acccf6e9337f65178bf

Observation b7b7176e-bea8-4b06-9999-8cea01f96982 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.057624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.057624Z digest=sha256:303a20f01bf6e59d5aac70c3ea6d116178744faaa49af8c5da26a9b3240715ec

Observation 7f637c6f-b367-4b98-bea7-b68c8f2122a3 · outbound

This paper cites Non-local neural networks.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Non-local neural networks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.063230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.063230Z digest=sha256:5690d55f5fbfa6192fe1ffb64a987adbd2f6205a47fa1284d017b1340b95b106

Observation 0374e256-6842-4722-ae9b-0d97a4800603 · outbound

This paper cites SOLO: Segmenting objects by locations.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SOLO: Segmenting objects by locations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.956507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.068431Z digest=sha256:baf4f34fcfe16d6c7aacfb6701aab18dac5ebc14356bce6e78a475b9159bab75

Observation 97186277-4ff4-47fe-a2ae-70c769f44316 · outbound

This paper cites Solov2: Dynamic and fast instance segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Solov2: Dynamic and fast instance segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.939830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.073957Z digest=sha256:e30b368f8e6da7de39a6d414d0108a0ae9f1d63d9187197e584db7d3bf8e3fd9

Observation df85ce4d-d2c9-4611-9ebd-2dc1a549d400 · outbound

This paper cites Images Speak in Images: A Generalist Painter for In-Context Visual Learning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Images Speak in Images: A Generalist Painter for In-Context Visual Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.080050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.080050Z digest=sha256:f1252836b5b430ef53842b52a6547a2e722c5d029a23400a7b2f44c83c5a2141

Observation 125b212e-419c-4753-8ec3-e38a42756858 · outbound

This paper cites SegGPT: Segmenting Everything In Context.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SegGPT: Segmenting Everything In Context

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.085777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.085777Z digest=sha256:70eae3ac9bd438a72bb2653bfcb32c9365af65512f883c8d101266afcdeb8368

Observation 59d13369-d98d-4855-b1fd-daaf17007b9c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Chain-of-thought prompting elicits reasoning in large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.922973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.091613Z digest=sha256:d4adbf41e6c938449c1dd55b198d1bf0dc5984b96dec3edce0d4590dda81d887

Observation 9ec67287-7a75-47a5-86a9-b6c218785426 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Gsva: Generalized segmentation via multimodal large language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.904733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.097842Z digest=sha256:e81cd762aec9ec1ee0e87ef8f0e60f492be593febfab3f0ee7887e6f9424f891

Observation 6104d1f6-685f-4683-bf9a-d92dec4c4a89 · outbound

This paper cites Alvarez, and Ping Luo.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Alvarez, and Ping Luo

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.104345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.104345Z digest=sha256:34ce0c232562532f60b82d010f6216928642f5048dad0f777c8e195f3ee43a14

Observation 4704578f-c2ad-4907-9140-7b5966b79f5d · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.109789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.109789Z digest=sha256:a74ffa372b3b58ba419794045d6bb2bfe65bd223d3020411cb18a4ecd042e6c2

Observation d1d07908-f8ea-431a-8854-4ab39d1af1b0 · outbound

This paper cites A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.859266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.115102Z digest=sha256:d0722eb92ef8440dd86ba92a67eb10d4a792827995397e44de589e224a079ad6

Observation 730d6885-8b49-43ac-81d3-67320b97a0be · outbound

This paper cites Fine-grained visual prompting, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Fine-grained visual prompting, 2023

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.840574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.121101Z digest=sha256:9567dc8e820e9f614e3ddea6bd0b95fce3179879bb44b98d5295bf8f749ba167

Observation 68db3912-cd72-4051-b9a4-11fd34e18873 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.126078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.126078Z digest=sha256:7dc2f93c7f8b49bdab4b7921a2133070ae076630c4b8ee2612e6949b1f5a9820

Observation 0a056f1c-adde-4864-8440-8f745078ba01 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.822743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.132233Z digest=sha256:df2121447039c542e489abc17309400f1d54aee612d6c3f09b2ba4ea340a9806

Observation 1cc26dc8-4858-4a84-acd7-69b225e49105 · outbound

This paper cites Osprey: Pixel un- derstanding with visual instruction tuning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Osprey: Pixel un- derstanding with visual instruction tuning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.805366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.138463Z digest=sha256:0603524804e444acecd47a2257b4af6ce1e92b203912177c97d25d8e581fdb13

Observation 548302c0-d01d-4a26-9eac-6ec52d771371 · outbound

This paper cites Sigmoid loss for language image pre-training,.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Sigmoid loss for language image pre-training,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.144523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.144523Z digest=sha256:f0a3d9602ef840956ef3ceec94a4e6377f939d16d3e6d836dd533554d8bd2456

Observation 7556df97-c4ee-4b7d-b8ab-f513b2f3bb51 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.151302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.151302Z digest=sha256:8ee5c8b595056c0eb417d0edea9af1fb3cede7bbf036845fc012cb22a400ccb5

Observation 5923443f-74a8-41ea-ae6e-c52518152f71 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.157888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.157888Z digest=sha256:912f2b5e25a056d2eade5250cf1a24e1c55898c72afe3d0de496a27fe5462d90

Observation 891d1b46-5fc9-4f41-a6f8-4fb9f8447cdb · outbound

This paper cites Groundhog: Grounding large language models to holistic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Groundhog: Grounding large language models to holistic segmentation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.775048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.163653Z digest=sha256:3d9cd6680e2901ceddcc8426b9db1a9a458560060cf5d99d16e78cedf880590c

Observation a43a1024-e009-476b-aede-3d65b8b1cb3a · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Psalm: Pixelwise segmentation with large multi-modal model

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.756026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.169375Z digest=sha256:1b23f5799080b65f05391fd901b7342e70cb1f6875420b0683ef091bf3f4ab66

Observation 46b773f9-a729-4c19-962d-315d1de1023b · outbound

This paper cites Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.739593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.174302Z digest=sha256:cce25ea19d938bff179424365cd3b977145d444e72bcb186ef06261d98db7ae7

Observation 6e715acd-3bbc-4c16-88b8-ecfbf0d317ba · outbound

This paper cites Scene parsing through ade20k dataset.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Scene parsing through ade20k dataset

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.722980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.179846Z digest=sha256:83994106d6f5978119f1bf0311390de77a58ffb73a0aec84cb07ad54232d0beb

Observation cf837357-61af-430d-ac08-29f3e48b1a29 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Seqtr: A simple yet universal network for visual grounding

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.704525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.185531Z digest=sha256:0f1b42027fc486f64647d79cc0cea383f9cf12a8e61cb81275cddd1ce5a067b0

Observation 25354d67-8d8e-4442-96ce-34acb6d4f0c4 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.190681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.190681Z digest=sha256:d938a51a63366384cd1d652ab3781826049196e91c4a1325a1089feb9174f786

Observation 9380f26d-e32a-4e8c-b96d-46447ed59b62 · outbound

This paper cites User: <IMAGE,MASK> Please segment target region with mask and corre- sponding category.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model User: <IMAGE,MASK> Please segment target region with mask and corre- sponding category

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.686647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:10:41.196244Z digest=sha256:7be009e04aa0f6ad4e0656d85cf3275435a1cdd9c5d05a8a784254e72c0e1b31

Pith citing papers

Observation 5a541e7d-6bdb-4f7b-8188-4377cd6f44ad · inbound

Bias from small-scale leakage in Pulsar Timing Array maps cites this paper.

Bias from small-scale leakage in Pulsar Timing Array maps ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:32.954256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:32.954256Z digest=sha256:6aec2975cbf114cca93994eda20f13c61aae16c53f114f69c2c1ffa00e3fba5c

Observation 115d2920-c1bf-478d-af38-9de4301a7d7e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.914750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:7baeb584cb318b31ed77612a810911724e48deafd9918991ba444f9019ff4d90