Pith. sign in

Paper Citation Record · LEDGER

DenseMLLM: Standard Multimodal LLMs for Dense Prediction

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2602.14134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.14134 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:23:35.957580Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T01:54:30.649092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.424891Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96f62d90-8e2e-41b6-ad2d-0f6fed872828 · outbound

This paper cites RLE string.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction RLE string

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.463345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.463345Z digest=sha256:a0ad07875fe34bb4142dcdf63e0b0b03a8d0798ac9f0cf5700904630ba06331e

Observation 658d868e-8250-4af4-8bc0-9ce006f3c349 · outbound

This paper cites an unresolved cited work.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.718786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.718786Z digest=sha256:c2f4835ecb62739246973da6c1d6eceb8f494cf3800c1b9a5df7c8c9aec86b68

Observation f46f50ed-9152-4b56-91ae-a312b9e0d20d · outbound

This paper cites S., and Lin, M.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction S., and Lin, M

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.302345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.302345Z digest=sha256:5cc5c40669fa0eec9cdd93f44a7e8182cd69f8e3fa5a5b1207f43225e77e7f2c

Observation a15d26ea-df99-4243-8734-283e6f38b8eb · outbound

This paper cites Swinmtl: A shared architecture for simultaneous depth estimation and se- mantic segmentation from monocular camera images.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Swinmtl: A shared architecture for simultaneous depth estimation and se- mantic segmentation from monocular camera images

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.473897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.473897Z digest=sha256:aa2136ffe5b99abbae65a0b1e92f3fbf5c80419a11e3d445826acf30feefa61f

Observation 75207e2f-7548-48ec-9a5b-1503e99d3947 · outbound

This paper cites Ufo: A unified approach to fine-grained visual perception via open-ended language interface.arXiv preprint arXiv:2503.01342, 2025a.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Ufo: A unified approach to fine-grained visual perception via open-ended language interface.arXiv preprint arXiv:2503.01342, 2025a

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.644056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.644056Z digest=sha256:2aa3db564d2e797def5891ac8c666eb05540156ba73ded9b5bc5e0fcc085ccd6

Observation 52b6dc11-bfab-4cda-ac4c-669226af52ec · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.805229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.805229Z digest=sha256:e5810b6e25b38a3c88370573f95a7af2fd390df5175b9758f09568e3b7b1d280

Observation bd6dfdfe-4afe-437f-a1a5-f4490fc688d6 · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.980496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.980496Z digest=sha256:264d814ed2bd7b0cb568e02667e3b3bfdeac20256a579acaf1d0411b0750ad70

Observation 6645d106-7b66-4803-9bb6-a7c6d0e446d4 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.129957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.129957Z digest=sha256:db7bb0c3df4ef98983eda5a40f2d9ef14010b97b6c93d239d4c4a0a44cf00a2e

Observation e52d6eb0-fe4a-4304-ad84-425e73b3fbdd · outbound

This paper cites Qwen2.5 Technical Report.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.241079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.241079Z digest=sha256:03acf002d94111427b021125c25464eb4b182f0c2c522dea9d94d00e8999754a

Observation 94644f63-9640-493d-983b-78c3f719e063 · outbound

This paper cites Visual representation alignment for multimodal large language models.arXiv preprint arXiv:2509.07979,.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Visual representation alignment for multimodal large language models.arXiv preprint arXiv:2509.07979,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.288627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.288627Z digest=sha256:c34416971064ab31ab8a2631a189d470400429e10f0d615481b69efd2ec05c62

Observation af58203f-071f-47fd-b286-8d4ed74fdd03 · outbound

This paper cites Semantic Segmentation for the Open World.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Semantic Segmentation for the Open World

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.634650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.634650Z digest=sha256:79814b1905c590f1ff63aaee8f9b0398e57303b2a7c94c50b9c76361db8ac7e8

Observation f7980be9-7a39-448f-b585-91aa30436ebd · outbound

This paper cites RLE string.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction RLE string

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.775714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.775714Z digest=sha256:9999307ab642bb3ede6609cbb80199e75208eb75ffa602e6f31fc7b5883502b8

Observation 721c078c-2fcd-468f-83c6-d0306fc8ebe2 · outbound

This paper cites During testing, the predicted values need to be de-quantized to obtain the real depth; otherwise, only relative depth is obtained.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction During testing, the predicted values need to be de-quantized to obtain the real depth; otherwise, only relative depth is obtained

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.957580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.957580Z digest=sha256:37566117e24cb46f75321ae72d1a02d38d1561c850e2c4014b150e5e0589289a

Observation 1e295596-64b2-4160-a1f5-c77d0116f190 · outbound

This paper cites During testing, we dequantize to the actual depths and exclude invalid depths.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction During testing, we dequantize to the actual depths and exclude invalid depths

Reference 1000

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.913684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.913684Z digest=sha256:3ab3617bc4b418863e01691a50fa00acae328304e64890fbed9b5a52889e117a

Observation e449bbf3-80e5-470f-b876-fd81be049334 · outbound

This paper cites Token activation map to visually explain multimodal llms.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Token activation map to visually explain multimodal llms

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.922033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.922033Z digest=sha256:e62ad5d81629e6f5225d3f453cd2b5c5d524b3cff9b81685fc900cf661ae4152

Observation 1451865a-7c9a-4791-9a06-15a32fff1dfe · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction DINOv2: Learning Robust Visual Features without Supervision

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.118285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.118285Z digest=sha256:c30b35352415c55b4f0e5c214ea07216b9e6ceb1bad0830864faeeb4eb2de27e

Observation 0872b2d1-4d96-4fd7-a576-8b93f4c585a0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.367262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.367262Z digest=sha256:748cccfd329797aeb22676d181c7f444db0772644a4494c8f3afefb9cf65e062

Observation 4008008e-1e56-4c2a-9d80-8aef88889994 · outbound

This paper cites an unresolved cited work.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Unresolved cited work

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.985256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.985256Z digest=sha256:32772aa15441feaee7002adcaa846d33ab38287d0a32febb7ec4f9b4fae37a13

Observation ac3b1f72-e304-4413-aa12-aa9b707e7c18 · outbound

This paper cites Depthlm: Metric depth from vision language models.arXiv preprint arXiv:2509.25413,.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Depthlm: Metric depth from vision language models.arXiv preprint arXiv:2509.25413,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.622893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.622893Z digest=sha256:ca58df99ba2a40a1e3332a10e9a6f1b21b98b3d344b4b93ce3b18ed9133afb5c

Observation 83e0c2c3-6ba3-4fb7-8596-b705792efcad · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.484278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.484278Z digest=sha256:e2561aed0c06d9256e16db414e5331aa2060f0c400bc22febc383fa5d99feb04

Observation 1e66a50c-6327-46c8-b119-652a6d6c6a69 · outbound

This paper cites UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.181055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.181055Z digest=sha256:68a6613d4d07b8d56fe13365061c7d9d97703327181c1970dcd21c251d70b38c

Observation 081f0df1-d124-4cc9-abe0-b3032db15574 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.793070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.793070Z digest=sha256:6517181f120b996b25bce6b0135269d7f2844b8f71d56fc0b975d80c306d1ff8

Observation c736002b-5762-44f3-92ec-012ce0ce4a4b · outbound

This paper cites Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K., Galley, M., and Gao, J.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K., Galley, M., and Gao, J

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.046942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.046942Z digest=sha256:77e1bb5c7ecbfb0c560276c40c964b22430ead73ea6fbf3a43650922b40f1596

Pith citing papers

Observation e2fbf0c8-a908-4dbb-a474-6a0934b4ef71 · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation DenseMLLM: Standard Multimodal LLMs for Dense Prediction

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.426147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:db5938008ed71d320737ccbb8308e85070ba77d6122ec76adc110660a7660c5b