Pith. sign in

Paper Citation Record · LEDGER

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

As of 23 August 2026, this Paper Citation Record lists 100 of 101 outbound references and 9 inbound Pith citation observations for arXiv:2506.24102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.24102 v1

Coverage vector

measured 100 of 101 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:40.777286Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:52:50.914242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:17.974489Z

Reference resolution

100 of 101 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a517616d-28d1-4ab3-ab06-c96b2e30ba7a · outbound

This paper cites Qwen2.5-VL Technical Report.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:37.747749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:37.747749Z digest=sha256:bbc048fd617e67344d8e090e52bc6720851b8bb3510ae4e66fb2c7b3ad86874b

Observation b7e8f173-a84d-4a68-85be-f207b761ac40 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:37.800973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:37.800973Z digest=sha256:5e84e7a24275f64f376b828631d666d1d852cdf641fe744af65a284c79775eeb

Observation b9aa99d3-ee94-4703-b2ef-5f43fe6a6ebc · outbound

This paper cites End-to-end object detection with transformers.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World End-to-end object detection with transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:37.884655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:37.884655Z digest=sha256:233ea2ae56bf385361e98e8bc71317f5f9578fd50df48713f0fdc3d7dced9c5f

Observation 0fc68726-06a3-4a0f-8de8-ede79d38e38f · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:37.952668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:37.952668Z digest=sha256:ec030a7ec3c35ef2e645d871365feeb6d9a1cb5f5eff97ebde29213ba11ecc19

Observation f09fa8bd-6a5a-4514-a61f-37094c21544b · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.026930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.026930Z digest=sha256:b32fde55c092c46868ba0608306ccb73c7088a02ec739645368441bcc88ce168

Observation 009dfa7e-fced-4aa2-b3a9-49c1447514c9 · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.IEEE TPAMI, 2017.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.IEEE TPAMI, 2017

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.100032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.100032Z digest=sha256:647db3c2c021265ef4fbf3dfe68bddf5bc483317d519fa49cfa4ec53a98f8fd2

Observation 2ab18dc2-767c-4897-a00b-ecbbf5059f1f · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Sharegpt4v: Improving large multi-modal models with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.168416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.168416Z digest=sha256:81e1deabc6c09f74fbdb672bd6dda554287b6740222cc36cd41ade0ce08c39ef

Observation 6fe10566-430e-4815-a496-2e863bbea96e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.248893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.248893Z digest=sha256:3677c2f779c2bdf8076936955037f68af560f5e720c2163b3f1afa1fc0e84cae

Observation a787fb3e-a4af-4da6-b1a4-471fc396660f · outbound

This paper cites Mmict: Boosting multi-modal fine-tuning with in-context examples.ACM TOMM, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mmict: Boosting multi-modal fine-tuning with in-context examples.ACM TOMM, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.314179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.314179Z digest=sha256:9b6702a426f500765e2817f6d8285eb52163cf796d5ff80ff092c8d98ce26c69

Observation 00641be1-9805-483a-a093-d7e1e84f9ff2 · outbound

This paper cites A generalist framework for panoptic segmentation of images and videos.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World A generalist framework for panoptic segmentation of images and videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.387930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.387930Z digest=sha256:2aed34b92bfd132772dc28f2dccff9fc42889a149cfba996496b68bbbd536e3f

Observation 97586feb-2cb6-4163-8f29-27f8764a0ea6 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.511775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.511775Z digest=sha256:eae03e72a73705a5ec663d946a2ff5f12851479a74158b2ec29ea54da0b47f36

Observation 6b4d502f-f185-45bb-8f7b-2f7b04494c20 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.683907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.683907Z digest=sha256:e8cad5754a4e2df9ace6d081ebfc387bb3f849832a9ff3b8a127cce3a364a7a9

Observation e7a9f7ce-d491-4732-9a9f-6a335dee48c2 · outbound

This paper cites Lmdeploy: A toolkit for compressing, deploying, and serving llm.https://github.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Lmdeploy: A toolkit for compressing, deploying, and serving llm.https://github

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.755578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.755578Z digest=sha256:f72f539aec7a34e4c7d32139def47e2675241078c6852961819bac6f1b7f2aaf

Observation 41ba3a1d-d81a-4a17-a5a4-0e43e0e5f766 · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.810255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.810255Z digest=sha256:09d82efbd632a62042a487c14a2d889f76a81fcfb5e6e825fdb7715ce47840b5

Observation 62fa214b-a681-4449-9fb1-8fa7716e92f9 · outbound

This paper cites COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:41.419910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:38.873138Z digest=sha256:8e787a54eeccb1eac5538fc5d01cc47c1fc93af46f94ea88ba1c7bf2dc90e27c

Observation fd30797a-9400-4a4f-a527-d611d4e00e20 · outbound

This paper cites Open-vocabulary universal image segmentation with maskclip.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Open-vocabulary universal image segmentation with maskclip

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.919524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.919524Z digest=sha256:026037d374865eb94e27d9ed9114701ba565f28217a32ad9d333b8077e53838c

Observation 16b4187a-f9f2-4dd2-acd0-843a8ed45312 · outbound

This paper cites On path to multimodal generalist: General-level and general-bench.ICML, 2025.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World On path to multimodal generalist: General-level and general-bench.ICML, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.971631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.971631Z digest=sha256:88e68021d6c97469a69f65d775ff608a194a9ea8353aae8248980779eb77985c

Observation 251abd9c-bc7b-485b-8c55-f4f2ad9bea19 · outbound

This paper cites Align-kd: Distilling cross-modal alignment knowledge for mobile vision-language model.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Align-kd: Distilling cross-modal alignment knowledge for mobile vision-language model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.024891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.024891Z digest=sha256:db48c70177b5e0edd1e67c8c70ef163647911cc0fbacaa67783111a8e52d1c51

Observation 51ce1e35-c333-4395-a02f-8310e9765c30 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.071048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.071048Z digest=sha256:866418769b0afbad313d7085b99a6e33ba13543b9697b7929a0d3874cce1c2c5

Observation a76d1f1c-6fda-47af-ab42-ef74e8a250c0 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.153550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.153550Z digest=sha256:b7a3c8579f507118608d1e50f64abf783d3255c5248c832c9a815fcc3c55f41b

Observation 865be4a2-489c-47a7-8a66-36dddf56bc2b · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.220822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.220822Z digest=sha256:88dd1a4e22758ef1b228f23181ea0d22bd7aba2ada1ce5b88a9307307db28561

Observation 016a5aa5-8bed-4c26-a7d3-c2f018e646f2 · outbound

This paper cites Fast R-CNN.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Fast R-CNN

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.286090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.286090Z digest=sha256:db5fa9dcc528d8609030a2139c67d987267a78a38fc47b970cbe770541758331

Observation 4a6c5b47-e72a-42bd-911d-b726c8c2e9cd · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.353274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.353274Z digest=sha256:ff99991312d85a26f707f4675efa5dc2a7a91ddde29d9d869941a2aca80d17a3

Observation 8045657a-32af-4ee3-8888-4b699826fde3 · outbound

This paper cites Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.419795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.419795Z digest=sha256:7452b04073abf5ebc9c46dccb70f4cf27256bdd9d2559be4c4315e04554bf178

Observation bad9970e-702d-482e-8bd4-74b851201b9e · outbound

This paper cites Mask R-CNN.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mask R-CNN

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.498070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.498070Z digest=sha256:8e82f602b0322179c432e08c8877c6beee38611b0e6b543b27619a0d76ea4664

Observation 93d3db71-0e2f-460b-82bb-417781bf135c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Lora: Low-rank adaptation of large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.554850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.554850Z digest=sha256:e7c8c3ddc657847d7e0837296f57d72ffcdb9c0a393117c96348c526747a7e57

Observation d8d231e1-25da-4d21-9123-7e2fae3108f1 · outbound

This paper cites GPT-4o System Card.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World GPT-4o System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.601377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.601377Z digest=sha256:adb6fd16a6062b6c176b738d4d61f9000869b64774d7909c8cf3951e28c52e6a

Observation 50c4069f-c7d7-4102-b753-1fef9742dde0 · outbound

This paper cites A diagram is worth a dozen images.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World A diagram is worth a dozen images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.661296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.661296Z digest=sha256:79a417978aeb27539caf7844fbfcec78de7f330ae81c3de27af8181838aa0411

Observation 29c7012d-5a25-432e-a20d-80f9c0efde03 · outbound

This paper cites Panoptic segmentation.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Panoptic segmentation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.728999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.728999Z digest=sha256:0f240bb2f988f8069022dd67a24dafdd2c45d89ae487efa43caa7eaeb378425d

Observation 534c5c57-4898-43da-9075-8315e1e35260 · outbound

This paper cites Segment anything.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Segment anything

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.781933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.781933Z digest=sha256:2a61a91223883ea96b977d457f4181fdf0bcb7a79ecf7055e6ad29bb80645400

Observation 68249ae9-326c-4e67-9b8c-99e4b686bc31 · outbound

This paper cites Dense-captioning events in videos.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Dense-captioning events in videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.846250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.846250Z digest=sha256:6eb3d49ba105c550848d1c090ad75f0fc88df383e78816e977b96c0047b4304a

Observation df4b7773-d011-4f2b-acd7-0cb8d9d5bd2f · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.874525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.874525Z digest=sha256:c573b20e8532aade16b5a2041093f356f1ce7229664780cd62782db3e501c10e

Observation f166d36b-6c5d-49bc-85eb-02739e935b59 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Lisa: Reasoning segmentation via large language model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.939420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.939420Z digest=sha256:15680301dd2530ba0e14313155d75cbafa39aca5f028bf90c7bfc6e61c069750

Observation c4d373a3-5f0d-4c36-8722-ff613582c522 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.999572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.999572Z digest=sha256:133c6e2701ac53ccf861e9aed97a2a8159af7e85cee6abea2df574c3444261ec

Observation 9a1bbaa1-b725-4fdb-8ac9-56aa74993404 · outbound

This paper cites Semantic flow for fast and accurate scene parsing.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Semantic flow for fast and accurate scene parsing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.120464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.120464Z digest=sha256:f6b692c62fa2906f5fc2e9fb8db99ef719e849d6271d7eb958d05c72b42e44c1

Observation 8dff9c5f-e137-4531-b6e8-dea61f2d7745 · outbound

This paper cites Tube-link: A flexible cross tube baseline for universal video segmentation.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Tube-link: A flexible cross tube baseline for universal video segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:46.573229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.248247Z digest=sha256:ce5adff6f561002f57ed7e556a5c890cdf49c50080f48c3b8d14cb0f18f9cb8b

Observation 3ef0665b-ba69-463a-a227-c9cadcb9cd18 · outbound

This paper cites Transformer-based visual segmentation: A survey.IEEE TPAMI, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Transformer-based visual segmentation: A survey.IEEE TPAMI, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:46.408374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.343543Z digest=sha256:b950639c9d41d41ba3009c95f714f7184d311373fca4f9b8249d3b31d6990ecf

Observation 2bbb1d49-443e-4f12-b427-d8766afb237b · outbound

This paper cites Panopticpartformer++: A unified and decoupled view for panoptic part segmentation.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Panopticpartformer++: A unified and decoupled view for panoptic part segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:46.251419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.396058Z digest=sha256:494c930f60fa60303b974364fa6a109f7d158514c5db1cc1e8ab61e7e2707577

Observation aa89047d-f226-4218-ab42-aecc1f300136 · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? InCVPR, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Omg-seg: Is one model good enough for all segmentation? InCVPR, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:46.132288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.510530Z digest=sha256:c5fe0d8b6dd4246b229d99f0e8d078a4ca38cef4b9f3bfaabcc3758ce449b98e

Observation 527de3c6-b531-4162-927d-0deac90836d7 · outbound

This paper cites Densefusion-1m: Merging vision experts for comprehensive multimodal perception.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Densefusion-1m: Merging vision experts for comprehensive multimodal perception

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:45.984022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.569012Z digest=sha256:098699290215620b39aa187b3cedd73fb8ee15a6bb7e2b039ab88d10c4f7ca69

Observation c0790401-28cc-4c64-868e-780d1a674f08 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.572938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.572938Z digest=sha256:fe6b063eb6d82d8c8faff0d8d85f2df50e306df777e4d51a7d36dbd1ba4eb3a8

Observation c844c7a3-b105-4555-b89b-bf055ebdeb44 · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.576534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.576534Z digest=sha256:947259ae6dac358667dd8c41e18e5c05269667554cd1f63a7ea32d51224efe4b

Observation 598a04af-4759-4cd5-ae13-94be05182c7a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.580439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.580439Z digest=sha256:85d7aa8f9d1802daa8b7c36ee7b01212debd8b32d83af9898604a017bc801bd8

Observation 88536864-af88-40b0-8df6-7f58519084cf · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.583636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.583636Z digest=sha256:f7a0234a4a67e91ee92bc3528ee2d47349998796ad9db89e6765342d2f657a4d

Observation c9ef35c5-618f-4079-a2a6-918379c65b15 · outbound

This paper cites Visual instruction tuning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Visual instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.586932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.586932Z digest=sha256:c3ca5a9878cf4ef3b603f399bdf5f61cb8e9918df33a657c403263b816a278fe

Observation f685c171-351f-4598-b460-061d8ea477a0 · outbound

This paper cites Improved baselines with visual instruction tuning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Improved baselines with visual instruction tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.590167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.590167Z digest=sha256:43074dec081d1f87a8942a9ae8b6e46b43458acb2b8ecdaa957e1027f89b443d

Observation b822752f-7f15-4cca-b248-8149c1ddf45a · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.593218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.593218Z digest=sha256:d9082935e2ee9b20accdf4d25fde844806674cc7d03497177fa53433e5bc16e8

Observation 9e1ecbfb-b11a-4f10-8b8e-68a4c107806e · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.596482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.596482Z digest=sha256:254be381b48da0a8ba2d3c90ad1e42af7c2e98ce6b4b1dcc161437ea923f5cf9

Observation 51407231-9259-4f1f-9ae4-19ccd0cb72e1 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mmbench: Is your multi-modal model an all-around player? InECCV, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.599775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.599775Z digest=sha256:bb93a016c7a3af7aa994000871d21ef7e91f86f6404da87342122587cb056031

Observation 32d988f5-3633-4586-a7d3-aa50531bf0c3 · outbound

This paper cites Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:45.818256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.602996Z digest=sha256:08eb7e86b07f633ae277e5bca57b69d3db6f2a6f52a58d3fca05f745f0554d20

Observation 771b976f-f12b-4fb9-b996-ba86b6d8470f · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.606698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.606698Z digest=sha256:634e16380ebbaea533925bfc7d0a36ae1fbf6fa1767277ea4bd96f71f3e14664

Observation 2316f26c-859e-4275-9f75-c4064a8bede1 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.609983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.609983Z digest=sha256:82dddf2e88eb6059372e5c29866d5b3903685ec32b3b45729377251fba5802c1

Observation 5526a603-514c-4ea4-8fc1-32e5e4e3c89e · outbound

This paper cites MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.613286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.613286Z digest=sha256:76ba2427f72fbbdc6f33c6c375ede5d4ecaf25150ae761ab60c77d5d35e523fa

Observation 9ac68ac8-901b-4519-9d09-e75b9bcd177f · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Generation and comprehension of unambiguous object descriptions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:45.643071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.616835Z digest=sha256:f2abd09499e0e9fcb23fe5c1b128d29f0345c66bd9886574fea0746d24eeba72

Observation f7a05ce1-34a3-48ad-b0ee-67f2aa27cd0e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Docvqa: A dataset for vqa on document images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.619942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.619942Z digest=sha256:2bb6a89a52f2dea1e43e0ba1b353812eea6b8d153d33014bfae16d814544dec4

Observation c008511e-efbf-4603-8c3e-76703c2a874e · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World DOCCI: Descriptions of Connected and Contrasting Images

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.623070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.623070Z digest=sha256:8de37f8bf4d2d26344b1ba68858e4a45c321c451c0ffde5a868e21373a7fb97a

Observation 1fd29755-a279-441d-b6b5-70abd5885946 · outbound

This paper cites Open world entity segmentation.TPAMI, 2022.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Open world entity segmentation.TPAMI, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:45.489298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.626880Z digest=sha256:b69f95c9a75fb3f1484faa8a5a84d28501b579cf86a98b6e5f9a9a582d1d486b

Observation 5c6244ae-b38c-4a04-977c-6d7226504cb9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Learning transferable visual models from natural language supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.630136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.630136Z digest=sha256:ac9d6fe511a5e22609bb9de3b9553e871c23d70b19b8aadae6cdb2c055506274

Observation 7305e160-f785-41c6-8060-27508a49775c · outbound

This paper cites Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.633432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.633432Z digest=sha256:acad2527e75493578753a32335b663ebb4f8eabd2836c8ac67f7518b45bce1df

Observation 3ba85684-9f45-4328-81b6-5933120051b0 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Glamm: Pixel grounding large multimodal model

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:45.259646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.637224Z digest=sha256:42589628a328d01f4b2d97eba8f07bd44899e2be4a012804da94ac358e07e106

Observation f255b326-5382-41ae-9632-a584e9951042 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixellm: Pixel reasoning with large multimodal model

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:45.091819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.640628Z digest=sha256:40b4283fa8641af4f7e132d8a63f7d45e2a9725e358ddcb06e171d7af7d80df3

Observation b36f3223-e009-45d6-862a-7cb8b8c7cfd0 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.NeurIPS, 2022.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Laion-5b: An open large-scale dataset for training next generation image-text models.NeurIPS, 2022

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.643719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.643719Z digest=sha256:97571d8c77708f0de2f74aab0bcacd279420156f8f97eb3eee719deee8654ba5

Observation 954b9cfb-a9d5-45a8-873d-1cb0b1de4485 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Objects365: A large-scale, high-quality dataset for object detection

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:44.894015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.647008Z digest=sha256:d42967ef8eb4d987496a8877f2a51c62a78d06f3ead1844a33c853311f092bbf

Observation c0c64e1e-61c4-407c-8a01-d765ca784027 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:44.714371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.650343Z digest=sha256:8ce79ca1521d289d1ced03934e156b1f5af56d906a0c54e7e1676c9be3d06180

Observation c589784d-a8e1-4873-abaf-3c7373c846d9 · outbound

This paper cites Aligning and prompting everything all at once for universal visual perception.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Aligning and prompting everything all at once for universal visual perception

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:44.503117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.653930Z digest=sha256:67d9b529eb80a436ce35563a725a06234b8beab25f70212f3f684971dead6ddf

Observation 506c686c-fa23-4a1e-ab93-1d3509cf6f53 · outbound

This paper cites Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.657128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.657128Z digest=sha256:4875360ae7a19863f960df3de6466990a4570ff3da1dab8879e9db78f67c831d

Observation f811175c-171c-4137-a96b-46026a58ac04 · outbound

This paper cites Seed1.5-VL Technical Report.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Seed1.5-VL Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.663603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.663603Z digest=sha256:18ada074b7d4f6f89df4322a4d6a9d59b8454a4be4feacf6bc86b466159d74b8

Observation d67485ec-2934-439e-ac44-936fa809d5ec · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Gemini: A Family of Highly Capable Multimodal Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.667140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.667140Z digest=sha256:00f7c8c06b9f95b191c2e055b8bd5ca3a5920feb91b81f93caae0746668790cf

Observation f41b03d3-62f3-4a2d-8e5e-8a54197c346f · outbound

This paper cites Yfcc100m: The new data in multimedia research.Communications of the ACM, 59(2):64–73, 2016.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Yfcc100m: The new data in multimedia research.Communications of the ACM, 59(2):64–73, 2016

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.670511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.670511Z digest=sha256:2196de006de099334dc7510ac511ab2a42e670fc9ff4a7a1ee47474e19882d5f

Observation 8ef410e2-0490-43a0-84eb-156ff7ce61d4 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.673570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.673570Z digest=sha256:0687a8aa87193f2ac0aa51f5f151d4c98cedd5526e10d8c5c667a60ff0694927

Observation 0680ee7b-5ee3-4dca-bb37-4758a8703e75 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.676960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.676960Z digest=sha256:e02616565ce0668efaadaa671ce00db332bea1c5714dfdbfef652aa7f2108bd6

Observation 16724b0e-b87a-4a6e-8d44-48b746bfe395 · outbound

This paper cites World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.680175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.680175Z digest=sha256:8edc1bf8429e6fc4517ad0ba4251e83752a2a11cd876d94c9816c2651a097955

Observation 81f90f83-af06-4238-a2e8-224eff335abd · outbound

This paper cites VGR: Visual Grounded Reasoning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World VGR: Visual Grounded Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.684250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.684250Z digest=sha256:d993b0f37cb43eee3bfc01de45749391485a37f2ac702f94d1dd1aed36752c8d

Observation 2bf8c0af-c9e1-4072-8f0b-704c64285cda · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World V3det: Vast vocabulary visual detection dataset

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:44.298107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.687959Z digest=sha256:f52ea8787b899403c7d0c47192291eb8d5d4b26fe0240d9ad18b385923134b20

Observation 647a6b22-6125-4f2f-b0af-167ba22b114d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.690961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.690961Z digest=sha256:dcb3da87a125726a20e2fbadeae8d69870b2b46e953d24085169eed5be27cdd5

Observation 82780df1-cb0c-4b10-a07a-133291ee74c9 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.694416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.694416Z digest=sha256:86ec9a4aa9c55596f5ab09b87c70093e2c15a157b58633dc5855beb1557f1c35

Observation aba661de-d16a-40ae-87d7-6d2c92714a83 · outbound

This paper cites The all-seeing project v2: Towards general relation comprehension of the open world.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World The all-seeing project v2: Towards general relation comprehension of the open world

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:44.109470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.697874Z digest=sha256:9f3daa2412165e9a554b5fbfb5f771b0807f38ec7f5b25bd1223ad11fa9cd0f7

Observation 88921f66-93cb-4832-9e49-5dff9ba3d940 · outbound

This paper cites Images speak in images: A generalist painter for in-context visual learning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Images speak in images: A generalist painter for in-context visual learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:43.872751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.701034Z digest=sha256:49c74f84864b45aaaf64808a546e579ac11d33065f2ca2ac73f4043ce5da9ee1

Observation e9caffc3-cc1a-43ad-8d9d-3f12d6fd71fc · outbound

This paper cites Controlmllm: Training-free visual prompt learning for multimodal large language models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Controlmllm: Training-free visual prompt learning for multimodal large language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:43.657982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.704032Z digest=sha256:0a876aae4f0df0b73296c77ba2cae11d102136be13fd76fb1672fd0b1bfc56a8

Observation b15ec7bb-0d6d-41d0-8242-2680524c2538 · outbound

This paper cites Clipself: Vision transformer distills itself for open-vocabulary dense prediction.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Clipself: Vision transformer distills itself for open-vocabulary dense prediction

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:43.454414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.707391Z digest=sha256:de46d8e769cc72354ed6363c46eee46681e294607481478455df88578c8cf6eb

Observation 72ca1fa0-bb09-431c-9aa5-d412737e49e9 · outbound

This paper cites Rap-sam:towards real-time all-purpose segment anything.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Rap-sam:towards real-time all-purpose segment anything

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:43.238441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.710740Z digest=sha256:f90534910116bcae5ad724e24885d77c2676906c017e5d4ea5a50d1402ead8cc

Observation a98feb04-96d7-46e7-a7c4-3a05c5ef5e00 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Visa: Reasoning video object segmentation via large language models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.713811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.713811Z digest=sha256:26a75dad7b1d591cdd9104aefee71a4ffc3442839ba9475fb03f8382fac72b4f

Observation dd6effe2-2bfd-4abc-9883-c49562682499 · outbound

This paper cites Qwen3 Technical Report.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Qwen3 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.717044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.717044Z digest=sha256:329bcaaae63863df0c5dfc0915ab6ce62d836fa68327ff794e0a9a322988ba74

Observation 78b64d86-590e-4d39-80fc-683b7a182919 · outbound

This paper cites Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:40.909214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.720476Z digest=sha256:35bada04b7031ca28140559777fa5ba12ef1d8e77027c23456cc7c87263b07e4

Observation 84d75ca1-68f0-4293-8cbe-8421400fde0e · outbound

This paper cites Modeling context in referring expressions.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Modeling context in referring expressions

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.723869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.723869Z digest=sha256:11b240c27e20d114edc53a19778ea39a6b465ae55111eea6a1040781e6de4eae

Observation 2dd084ac-2f01-4ffe-b665-98d5ac0b23c5 · outbound

This paper cites Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:42.948227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.727389Z digest=sha256:315469303d9231b775acd2cd98b47d31de539c06aa6e1039f7b586d4bbf8c00e

Observation 326cc963-2d20-4580-ae9d-90299834ae37 · outbound

This paper cites Open-vocabulary sam: Segment and recognize twenty-thousand classes interactively.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Open-vocabulary sam: Segment and recognize twenty-thousand classes interactively

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.730926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.730926Z digest=sha256:51f03bf1a61b55d802184d800d506c673409524d236fd4e6001ed410c4596792

Observation 91968c22-12af-4077-a3ca-2209fb457c4e · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.734553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.734553Z digest=sha256:44ddd6ba31c67368035e8251ec33e9e2a76c71eb5c73c873d48594f9656d1143

Observation b4f3a023-64e9-4de3-af38-b4f4e40c4ee9 · outbound

This paper cites Instruction-guided multi-granularity segmentation and captioning with large multimodal model.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Instruction-guided multi-granularity segmentation and captioning with large multimodal model

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:42.738685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.738303Z digest=sha256:49110512b8dffac0c81754f5f56a300cfe1c96c274e6e9d31e22e15caca109cc

Observation 4217cfe9-f7c7-46c7-afd9-3e655ee88f6a · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Osprey: Pixel understanding with visual instruction tuning

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:42.489058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.741516Z digest=sha256:11c2561de073dd778bb376bd1e42966847c1cc9ac0c691725bf6f1cc5fbde0b4

Observation 6584a801-5585-4a28-945b-44b6f37e66e2 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.744822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.744822Z digest=sha256:845683537df0298b25e91264602563f53310701886a6c67685b89cddbb294351

Observation 975c5b05-e79a-4591-9cf7-cb0d3a5778e9 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:42.234512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.748210Z digest=sha256:509f18cd928f7bff196a2cb1742c893ca106a609f8ff1b4b91650c785857bef1

Observation 925c7ed3-2ea6-488a-9936-4a04bdc85585 · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.751792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.751792Z digest=sha256:440c863c042f6ec10876035fe92fb705d630dd35cea0bb77b56af66dff9a7ecb

Observation 85e6f998-a196-4e31-9939-9d1029a12a5b · outbound

This paper cites Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.755138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.755138Z digest=sha256:3963b1fe374ed34e377e7b26eec71226dc9ef7af1e8456c7353ff2b705a11c3a

Observation 1e69aa0b-f4c7-4968-b170-7e020521aa24 · outbound

This paper cites Enhancing multimodal large language models complex reason via similarity computation.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Enhancing multimodal large language models complex reason via similarity computation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:41.966261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:27:40.758700Z digest=sha256:02d59a0d406fef7cead92f1e869706e6e32b695cce72979be5f5aa6e34685cc1

Observation 176d8e49-9d01-4eae-b0d3-8f678ab40867 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.762326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.762326Z digest=sha256:ff04964fdc999c73a53799550876639ebea113db77fd00c368f7d5cfec09c635

Observation 3f7e994b-b245-48c3-bdd9-998ef8bb5e91 · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.766319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.766319Z digest=sha256:5e53d3718aa3de05b36505d9d4d2393f6d5b1f1736098dbfac72ae26b891cf25

Observation 21a32f5d-15ab-497f-9a92-3789a304f80c · outbound

This paper cites Recognize anything: A strong image tagging model.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Recognize anything: A strong image tagging model

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.769992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.769992Z digest=sha256:514654dd88c3ee7549a6b250d078402b1d571500c89174888e10c1fe58c4b120

Observation 004cad2c-8fb4-4b9f-9cd6-db5b66514477 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MLVU: Benchmarking Multi-task Long Video Understanding

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.773704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.773704Z digest=sha256:98cd713b076c73262dc0cccd6e62efc5cde092a3e781d7f0be1ed7417b371c61

Observation 8878ef5f-2876-42fe-ae8d-403b1ca08e06 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.777286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.777286Z digest=sha256:e96a8dfbb261e60d9fc55a9076a0af0117ddc761b2654d5d0df0e4efd65353df

Pith citing papers

Observation 6acbc47d-b25b-403a-95df-24f4be8505bb · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.054292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.054292Z digest=sha256:a5b932d23285d648602fbb7640035da366254be2f64a46a31d5d86f30c9346ec

Observation 7ca2ec95-3991-4fe6-8abb-0740c4141272 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.818295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.818295Z digest=sha256:42aad4002990555eaabfb8fae966fc07825ef636f722b92b497856648ec5c36c

Observation 4b910caf-778d-4e25-ab97-dccd7d3bea35 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:07.535428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:e45115175bb7bbf3025101b52e74abff3386ad48c0ccd64c986b3d797c218da4

Observation df075b50-a630-4695-907b-0e677b6d790a · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.757959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:839df21d9d7178d755460aefc55ad5f6fca0c4d94220f88577e582029ec53f0e

Observation 6d5d9761-0f57-4a1e-915f-e3d09818a77e · inbound

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models cites this paper.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.978718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:d6dc7406e44c8776905104e7d8aeeb67855b0b81eaf02779266f22c793760786

Observation 221ab060-7b89-4d7f-9c15-25a93ec708b8 · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.176750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:05711f9a0135f7397b0c46d29ef6065f11e84e5690c0c3eebb2084e57cd8eb7f

Observation dadaec08-c2f2-4c26-924b-69885e6bc720 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:49b1e259e5f902959522ff6a24cb7f3a65ed32b6f85ba651fd133842d5f144c0

Observation 08ef701c-714a-4f5b-ae1e-6708a30a4cab · inbound

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering cites this paper.

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:52:50.914242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:52:50.914242Z digest=sha256:400b8e8642701b2e0d0e3f5c22ce39d1286f2fbca90c5e7827447a43384dff56

Observation 89047adb-f0dd-4db6-8626-d5c982287beb · inbound

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment cites this paper.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.366022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.366022Z digest=sha256:2a8fce979f09c1008f1843f3ae98f13fb089de4127bd81121ef97b1c927d1d73