Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:40.777286Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 101 outbound references and 9 inbound Pith citation observations for arXiv:2506.24102.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:40.777286Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:52:50.914242Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:49:17.974489Z
100 of 101 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a517616d-28d1-4ab3-ab06-c96b2e30ba7a · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e8f173-a84d-4a68-85be-f207b761ac40 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9aa99d3-ee94-4703-b2ef-5f43fe6a6ebc · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World End-to-end object detection with transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc68726-06a3-4a0f-8de8-ede79d38e38f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09fa8bd-6a5a-4514-a61f-37094c21544b · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009dfa7e-fced-4aa2-b3a9-49c1447514c9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.IEEE TPAMI, 2017
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab18dc2-767c-4897-a00b-ecbbf5059f1f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Sharegpt4v: Improving large multi-modal models with better captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe10566-430e-4815-a496-2e863bbea96e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a787fb3e-a4af-4da6-b1a4-471fc396660f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mmict: Boosting multi-modal fine-tuning with in-context examples.ACM TOMM, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00641be1-9805-483a-a093-d7e1e84f9ff2 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World A generalist framework for panoptic segmentation of images and videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97586feb-2cb6-4163-8f29-27f8764a0ea6 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b4d502f-f185-45bb-8f7b-2f7b04494c20 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7a9f7ce-d491-4732-9a9f-6a335dee48c2 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Lmdeploy: A toolkit for compressing, deploying, and serving llm.https://github
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ba3a1d-d81a-4a17-a5a4-0e43e0e5f766 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62fa214b-a681-4449-9fb1-8fa7716e92f9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fd30797a-9400-4a4f-a527-d611d4e00e20 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Open-vocabulary universal image segmentation with maskclip
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b4187a-f9f2-4dd2-acd0-843a8ed45312 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World On path to multimodal generalist: General-level and general-bench.ICML, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251abd9c-bc7b-485b-8c55-f4f2ad9bea19 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Align-kd: Distilling cross-modal alignment knowledge for mobile vision-language model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ce1e35-c333-4395-a02f-8310e9765c30 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76d1f1c-6fda-47af-ab42-ef74e8a250c0 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 865be4a2-489c-47a7-8a66-36dddf56bc2b · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016a5aa5-8bed-4c26-a7d3-c2f018e646f2 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Fast R-CNN
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6c5b47-e72a-42bd-911d-b726c8c2e9cd · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8045657a-32af-4ee3-8888-4b699826fde3 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bad9970e-702d-482e-8bd4-74b851201b9e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mask R-CNN
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d3db71-0e2f-460b-82bb-417781bf135c · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Lora: Low-rank adaptation of large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d231e1-25da-4d21-9123-7e2fae3108f1 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World GPT-4o System Card
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50c4069f-c7d7-4102-b753-1fef9742dde0 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World A diagram is worth a dozen images
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c7012d-5a25-432e-a20d-80f9c0efde03 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Panoptic segmentation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534c5c57-4898-43da-9075-8315e1e35260 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Segment anything
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68249ae9-326c-4e67-9b8c-99e4b686bc31 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Dense-captioning events in videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4b7773-d011-4f2b-acd7-0cb8d9d5bd2f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f166d36b-6c5d-49bc-85eb-02739e935b59 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Lisa: Reasoning segmentation via large language model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d373a3-5f0d-4c36-8722-ff613582c522 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1bbaa1-b725-4fdb-8ac9-56aa74993404 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Semantic flow for fast and accurate scene parsing
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dff9c5f-e137-4531-b6e8-dea61f2d7745 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Tube-link: A flexible cross tube baseline for universal video segmentation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3ef0665b-ba69-463a-a227-c9cadcb9cd18 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Transformer-based visual segmentation: A survey.IEEE TPAMI, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2bbb1d49-443e-4f12-b427-d8766afb237b · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Panopticpartformer++: A unified and decoupled view for panoptic part segmentation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa89047d-f226-4218-ab42-aecc1f300136 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Omg-seg: Is one model good enough for all segmentation? InCVPR, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 527de3c6-b531-4162-927d-0deac90836d7 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Densefusion-1m: Merging vision experts for comprehensive multimodal perception
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0790401-28cc-4c64-868e-780d1a674f08 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c844c7a3-b105-4555-b89b-bf055ebdeb44 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World RAIN: Your Language Models Can Align Themselves without Finetuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598a04af-4759-4cd5-ae13-94be05182c7a · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88536864-af88-40b0-8df6-7f58519084cf · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9ef35c5-618f-4079-a2a6-918379c65b15 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Visual instruction tuning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f685c171-351f-4598-b460-061d8ea477a0 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Improved baselines with visual instruction tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b822752f-7f15-4cca-b248-8149c1ddf45a · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1ecbfb-b11a-4f10-8b8e-68a4c107806e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51407231-9259-4f1f-9ae4-19ccd0cb72e1 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mmbench: Is your multi-modal model an all-around player? InECCV, 2024
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d988f5-3633-4586-a7d3-aa50531bf0c3 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 771b976f-f12b-4fb9-b996-ba86b6d8470f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2316f26c-859e-4275-9f75-c4064a8bede1 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5526a603-514c-4ea4-8fc1-32e5e4e3c89e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac68ac8-901b-4519-9d09-e75b9bcd177f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Generation and comprehension of unambiguous object descriptions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7a05ce1-34a3-48ad-b0ee-67f2aa27cd0e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Docvqa: A dataset for vqa on document images
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c008511e-efbf-4603-8c3e-76703c2a874e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World DOCCI: Descriptions of Connected and Contrasting Images
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fd29755-a279-441d-b6b5-70abd5885946 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Open world entity segmentation.TPAMI, 2022
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c6244ae-b38c-4a04-977c-6d7226504cb9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Learning transferable visual models from natural language supervision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7305e160-f785-41c6-8060-27508a49775c · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba85684-9f45-4328-81b6-5933120051b0 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Glamm: Pixel grounding large multimodal model
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f255b326-5382-41ae-9632-a584e9951042 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixellm: Pixel reasoning with large multimodal model
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b36f3223-e009-45d6-862a-7cb8b8c7cfd0 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Laion-5b: An open large-scale dataset for training next generation image-text models.NeurIPS, 2022
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954b9cfb-a9d5-45a8-873d-1cb0b1de4485 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Objects365: A large-scale, high-quality dataset for object detection
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0c64e1e-61c4-407c-8a01-d765ca784027 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c589784d-a8e1-4873-abaf-3c7373c846d9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Aligning and prompting everything all at once for universal visual perception
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 506c686c-fa23-4a1e-ab93-1d3509cf6f53 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f811175c-171c-4137-a96b-46026a58ac04 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Seed1.5-VL Technical Report
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67485ec-2934-439e-ac44-936fa809d5ec · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Gemini: A Family of Highly Capable Multimodal Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41b03d3-62f3-4a2d-8e5e-8a54197c346f · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Yfcc100m: The new data in multimedia research.Communications of the ACM, 59(2):64–73, 2016
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef410e2-0490-43a0-84eb-156ff7ce61d4 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0680ee7b-5ee3-4dca-bb37-4758a8703e75 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16724b0e-b87a-4a6e-8d44-48b746bfe395 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f90f83-af06-4238-a2e8-224eff335abd · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World VGR: Visual Grounded Reasoning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf8c0af-c9e1-4072-8f0b-704c64285cda · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World V3det: Vast vocabulary visual detection dataset
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 647a6b22-6125-4f2f-b0af-167ba22b114d · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82780df1-cb0c-4b10-a07a-133291ee74c9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba661de-d16a-40ae-87d7-6d2c92714a83 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World The all-seeing project v2: Towards general relation comprehension of the open world
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 88921f66-93cb-4832-9e49-5dff9ba3d940 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Images speak in images: A generalist painter for in-context visual learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9caffc3-cc1a-43ad-8d9d-3f12d6fd71fc · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Controlmllm: Training-free visual prompt learning for multimodal large language models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b15ec7bb-0d6d-41d0-8242-2680524c2538 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Clipself: Vision transformer distills itself for open-vocabulary dense prediction
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 72ca1fa0-bb09-431c-9aa5-d412737e49e9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Rap-sam:towards real-time all-purpose segment anything
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a98feb04-96d7-46e7-a7c4-3a05c5ef5e00 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Visa: Reasoning video object segmentation via large language models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6effe2-2bfd-4abc-9883-c49562682499 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Qwen3 Technical Report
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b64d86-590e-4d39-80fc-683b7a182919 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 84d75ca1-68f0-4293-8cbe-8421400fde0e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Modeling context in referring expressions
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd084ac-2f01-4ffe-b665-98d5ac0b23c5 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 326cc963-2d20-4580-ae9d-90299834ae37 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Open-vocabulary sam: Segment and recognize twenty-thousand classes interactively
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91968c22-12af-4077-a3ca-2209fb457c4e · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f3a023-64e9-4de3-af38-b4f4e40c4ee9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Instruction-guided multi-granularity segmentation and captioning with large multimodal model
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4217cfe9-f7c7-46c7-afd9-3e655ee88f6a · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Osprey: Pixel understanding with visual instruction tuning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6584a801-5585-4a28-945b-44b6f37e66e2 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 975c5b05-e79a-4591-9cf7-cb0d3a5778e9 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 925c7ed3-2ea6-488a-9936-4a04bdc85585 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e6f998-a196-4e31-9939-9d1029a12a5b · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e69aa0b-f4c7-4968-b170-7e020521aa24 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Enhancing multimodal large language models complex reason via similarity computation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 176d8e49-9d01-4eae-b0d3-8f678ab40867 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7e994b-b245-48c3-bdd9-998ef8bb5e91 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a32f5d-15ab-497f-9a92-3789a304f80c · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Recognize anything: A strong image tagging model
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 004cad2c-8fb4-4b9f-9cd6-db5b66514477 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MLVU: Benchmarking Multi-task Long Video Understanding
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8878ef5f-2876-42fe-ae8d-403b1ca08e06 · outbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6acbc47d-b25b-403a-95df-24f4be8505bb · inbound
Kwai Keye-VL Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca2ec95-3991-4fe6-8abb-0740c4141272 · inbound
Kwai Keye-VL 1.5 Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b910caf-778d-4e25-ab97-dccd7d3bea35 · inbound
Let ViT Speak: Generative Language-Image Pre-training DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df075b50-a630-4695-907b-0e677b6d790a · inbound
Let ViT Speak: Generative Language-Image Pre-training DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6d5d9761-0f57-4a1e-915f-e3d09818a77e · inbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 221ab060-7b89-4d7f-9c15-25a93ec708b8 · inbound
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dadaec08-c2f2-4c26-924b-69885e6bc720 · inbound
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ef701c-714a-4f5b-ae1e-6708a30a4cab · inbound
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89047adb-f0dd-4db6-8626-d5c982287beb · inbound
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.