Pith. sign in

Paper Citation Record · LEDGER

LMM-Det: Make Large Multimodal Models Excel in Object Detection

As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2507.18300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18300 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:41.385870Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58ed4f22-b4f8-40ad-a4d9-359276a652bb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.331412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.111541Z digest=sha256:8fb6f2e1dda4d330662f36504249db50d913491142e607f85067e6c4a9b3c8ad

Observation 27805ef5-c316-4dcb-8327-ef1aafee5b34 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.118189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.118189Z digest=sha256:748b9c92f0c360875317f8a4d7e1b844300acfe938174753419ea4f2961bd29d

Observation c03da974-5d73-4943-9c95-182c33f9a3a4 · outbound

This paper cites Cascade r-cnn: Delving into high quality object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Cascade r-cnn: Delving into high quality object detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.316592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.124440Z digest=sha256:2900cd46db981c89a8eb626a0a8e217fcda613708e869c35e2e438da4ba7b60f

Observation d6888f5f-6120-4b14-bb59-2471112cb795 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Honeybee: Locality-enhanced projector for multimodal llm

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.300835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.129824Z digest=sha256:c6b17cbfc748103256f65dcabb793f029fbea10fb10105cc903005bd135bc8bf

Observation 9d2577a7-ffdb-4c06-a5e5-63e52e641ca9 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

LMM-Det: Make Large Multimodal Models Excel in Object Detection X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.134807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.134807Z digest=sha256:219f12889c8dd0dc9fa94d8491527a59aa773a0affdda3b6454b54e126bcf944

Observation 5d46851c-81b4-46c9-a20b-822e1986d547 · outbound

This paper cites Allava: Harnessing gpt4v- synthesized data for a lite vision-language model, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Allava: Harnessing gpt4v- synthesized data for a lite vision-language model, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.285513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.140378Z digest=sha256:6db3f8bbde135425b373c8bef36ac2d37413bb23638d779835893264b43e4d16

Observation 5d0fbbc1-6be6-4921-b7d9-c4527c2acd88 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.145863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.145863Z digest=sha256:b9157d3c445726e1bfac966611a8106b9682ee7ae8e0ae437c79fc2567ea7536

Observation f2681c6b-841a-4a05-95db-0ff91859d98d · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.150411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.150411Z digest=sha256:e0c77c10e2f5b33c12fc0d0dccef9cccc816ba5e98269604deab2c94865e4898

Observation a80e0917-a2be-41b4-a030-a1668d3b4f21 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Glm: General language model pretraining with autoregressive blank infilling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.155155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.155155Z digest=sha256:b7cef84044eb4c7c2ac748b4675b596921600a29293df1edbb40fe8c50aa45f3

Observation d39d1c4f-789e-4370-bc1d-855a369de980 · outbound

This paper cites LVIS: A dataset for large vocabulary instance segmentation.

LMM-Det: Make Large Multimodal Models Excel in Object Detection LVIS: A dataset for large vocabulary instance segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.159873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.159873Z digest=sha256:5105b49a6ff7ffa12c18ee23adfad5e4551e97412d9e197f6bafe93e2c3b5dd3

Observation 344af88c-0b80-4af1-9c1c-219d6367009c · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Efficient Multimodal Learning from Data-centric Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.164283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.164283Z digest=sha256:5400412b722eb4c87f1ca50a5884c8625d15d38d39ba455b2821d58fe7f257e0

Observation 35d22764-04c0-456f-8ac3-6dd7db62aadc · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

LMM-Det: Make Large Multimodal Models Excel in Object Detection CogVLM2: Visual Language Models for Image and Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.169091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.169091Z digest=sha256:1b12a953a83c80618cb87f1ed665969f299292dd65323edc3e3f1adc0e41010a

Observation b80f73ab-759f-4a65-9d7a-ee91caaefc00 · outbound

This paper cites Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.249214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.173863Z digest=sha256:9e06c215ac4aa89c46baa3ef0e0eeb554ee65dc94811e1bd649c0719c97f7b7e

Observation 1929c7a6-2409-45b9-b92f-c3a2d481be66 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

LMM-Det: Make Large Multimodal Models Excel in Object Detection mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.232628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.178318Z digest=sha256:111b02055b288d4b0874987f981ccbd549b67779b48613251c968efb44249ccc

Observation 51dc8c80-68d7-4889-8fcd-79fc9d9758a3 · outbound

This paper cites Detrs with hybrid matching.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Detrs with hybrid matching

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.216107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.182911Z digest=sha256:5ff21420e306e2f8320014a5ed69a2f4d7ba73647bb4738ccfff9435ce06284c

Observation 9a0e7103-7d10-4fe2-9e0e-4f07eb61bc5c · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.187182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.187182Z digest=sha256:26eb9984088bd015029c342154eab7c3df5d536631006268f1f5d4de11bec794

Observation 9800eeba-e1b2-4715-a7a4-646b4b1401ad · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Openimages: A public dataset for large-scale multi-label and multi-class image classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.187371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.192354Z digest=sha256:beb7c561b8b67593870baf5b685967417e5f7319af9ebd2efa38bab471cfa2f7

Observation d39e0ca0-c515-4ee9-844a-d43e7d6201cd · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Building and better understanding vision-language models: insights and future directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.196489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.196489Z digest=sha256:972beaa47f8fe64cb2a87a706d60ca9d2e7a9a4951cdcf436faf7fbd9569f102

Observation 476c6df1-cf32-4b2d-aa5b-90eddafff098 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.171495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.201411Z digest=sha256:1d807c7ad40357130616f1977b225e00f68876f1415eeda37ef84106838ba826

Observation 565cd5ab-fd74-4391-99f2-b2b06f98ee1e · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Monkey: Image resolution and text label are important things for large multi-modal models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.154303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.205867Z digest=sha256:32002a5d14423160ec3f481e205fe9c0f80f1d6e987ac88c5f3166f6c552218d

Observation 98281367-7eab-4f27-84a6-8d5c990339cc · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Generative region-language pretraining for open-ended object detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.137555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.210434Z digest=sha256:10a929806d050ab6869add032f72d09760613af6783202558016fb4826cfafcf

Observation 082175c0-5c53-4bd6-8a3e-494d025f87be · outbound

This paper cites Microsoft coco: Common objects in context.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.121477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.214842Z digest=sha256:cee2b6823847f97e3ec66574d0fcacd31cfe28a971637de2dc58bb182f6f0368

Observation ee13eab0-2b8b-45fe-a68a-0c7b0b986ce9 · outbound

This paper cites Visual instruction tuning.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.103453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.219573Z digest=sha256:91d4c332d24468fc8b979f8527b7be1107b0a02e2d9d98f4d86e11c898523776

Observation 1b120a8b-eab7-47e7-bd7f-7b30a3a458b1 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.088137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.224159Z digest=sha256:d31f55a6217cb7a46ca3c3f66f0633fba98243a6c2b700b796ccce6cef4cb094

Observation e431bbaa-d72d-4f2a-8b53-69f2174f4a63 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.229099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.229099Z digest=sha256:2013c9def39486e9103c8fb22d6a8be31dce98efcb6771ce8960ceaf533993af

Observation 98756f12-295d-4452-a514-451e1932840d · outbound

This paper cites Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:20:41.658332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.233757Z digest=sha256:4c13bc8bf9d9daeb4a0dcf0ca1a89639fde979fce6983db8be4730facdc1c9e3

Observation c61919d6-a94b-425b-b8c5-06915cb95c05 · outbound

This paper cites Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.238757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.238757Z digest=sha256:6ef5e01cc936535401c65196ab74535734def9cca3881e06ea53109704cbe5e9

Observation 052f5567-27c8-4b4c-be71-d7197552d3df · outbound

This paper cites Scal- ing open-vocabulary object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Scal- ing open-vocabulary object detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.071999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.243496Z digest=sha256:c420db20ebe90a1a76ea90c7ad1a2faabffc525a09f7b206d5568f7cadafaa8d

Observation a00781fd-abed-48a2-acaa-09997f4b0545 · outbound

This paper cites an unresolved cited work.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:42.054615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.248336Z digest=sha256:b27df3a210df3007abe32290d6124de1b886a18b956e9759e8212eaf40b8e228

Observation 2f254c71-48af-45ae-932f-27a1aa2fa4ad · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.253210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.253210Z digest=sha256:60d722f479f370c8b71c88d48f243c3c412703826dbf761b9d0abbced8a050c8

Observation b5139c15-b74f-43f1-bee1-28ef7c17ab29 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Learning transferable visual models from natural language supervision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.039031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.258353Z digest=sha256:77c6e9321e175263052527f18f95cf8004fe5eada69b6b279bc3cc46a844b694

Observation 0b4a0aea-8d96-4d95-9530-be0dd34651e0 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.262855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.262855Z digest=sha256:18e3d5d5ca40bea03274548328a914b3907b1ab2d4443f9382a34eff4dad8b53

Observation 837f13fa-8959-442a-b801-bc624c644b0d · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.010266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.267864Z digest=sha256:d6a83bc6fb31754cf45a4854625c8b569cd64b6fc4033f73ae004df44dbcd115

Observation 6cd99452-18f4-4f55-99e5-0da38c2b3aa0 · outbound

This paper cites Grounding dino 1.5: Advance the "edge" of open-set object detection, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Grounding dino 1.5: Advance the "edge" of open-set object detection, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.993470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.272251Z digest=sha256:ee9ce77db803825acc1493ba02955a281d1cb0b75d1517b67db709144409d7f4

Observation 49ccb3ad-52a0-4065-a8cd-990bbf6c9447 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Objects365: A large-scale, high-quality dataset for object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.978046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.276858Z digest=sha256:f36b75834228cfa06126b21e5f352f877687c1242b5b25073e7dec80a49df1c7

Observation b870a08f-6937-4ea0-895f-4d4287904f6e · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.962458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.281336Z digest=sha256:4def294e13abbb46e2a5ad9aed0e950f14158c278a2958c8716467966e110f71

Observation 0322734e-0a1e-4d52-b725-2d6e3fae8019 · outbound

This paper cites Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.947188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.285832Z digest=sha256:6cf5a2456c3bbdfa25d1a99130ae7443b1b01856e533ef92b594412898bb1160

Observation 468d53aa-7fe4-4c97-8da8-343b9e703d7b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.290110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.290110Z digest=sha256:c9b324e4281ee23d8239e0371298cd9fddfe0ed372bea65c26f6ed5a27dc0bd2

Observation 8fb52397-f1fb-4d66-8565-f7bb61048fbf · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.295610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.295610Z digest=sha256:493a5ec763996b87b87c98eca032fcbbc0ac1d13dd2fa262a8253d62c73ba02c

Observation e79f9c03-8f8e-4163-8f48-d84cf427e9cc · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

LMM-Det: Make Large Multimodal Models Excel in Object Detection General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.300411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.300411Z digest=sha256:58e101e4e81958de71b93c22e8ec869dde6d8d6ff285ddd6316f1e2ffbbae7c1

Observation be62617d-1ea0-40d5-a7ab-2180920b719e · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Skywork: A More Open Bilingual Foundation Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.304607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.304607Z digest=sha256:9e4ca63478228bd630f3d694d74eab2fdee3abec7a890696d19d0c83f9ea3d2b

Observation ff4647ae-7bcd-46f1-b7d6-d1feb6494acd · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.931671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.309479Z digest=sha256:ab1db7ba616594e2f6ebf1b792958aa2df6cc68138570bdf4f95986f99f7c279

Observation 8ca45a51-e777-490b-9d47-ac8345db92a6 · outbound

This paper cites Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.915560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.315064Z digest=sha256:8b9a57d76b2f396a1d0fec3645daf10dc95276151d6a84b85d1231cb3a0d1027

Observation 046407b3-8617-4bad-93dd-d6babea75009 · outbound

This paper cites Ccmb: A large-scale chinese cross- modal benchmark.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ccmb: A large-scale chinese cross- modal benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.899307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.319711Z digest=sha256:2fd4efb883a4d84b93737cb302615479a330cc153027eb37004ce237e7c3cc7b

Observation 4c9c66b6-a221-4452-863c-7827fa9e2d26 · outbound

This paper cites FG-CLIP: Fine-Grained Visual and Textual Alignment.

LMM-Det: Make Large Multimodal Models Excel in Object Detection FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.324154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.324154Z digest=sha256:486d189aadc32dabf78e9fca154569fff95154fd1b9f2ee68b310529be1e62b6

Observation a4af59aa-d814-470f-b77b-9845ee14c2f8 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Llava-cot: Let vision language models reason step-by-step, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.882457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.328936Z digest=sha256:e0e243bfa28b3e7caa056eb9535c9a33de865ad12b89760247ebdc6db4b0e3d8

Observation e7e450db-e393-40cb-b8f8-ba493cc0aafc · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

LMM-Det: Make Large Multimodal Models Excel in Object Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.333364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.333364Z digest=sha256:4ae6bb083c6665400e9d2ea98cf773d12c29ec63227108e5dc3537a104c903af

Observation 36d87560-cc15-4da0-9f18-62bc53828daa · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

LMM-Det: Make Large Multimodal Models Excel in Object Detection mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.337989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.337989Z digest=sha256:f4db03c5ee7870fb5fc0dc8e1e31565130ab6b54ab24cbdc59d44200e3113ecd

Observation 25938f21-ddf7-4d62-b1a7-7c9a3aed47e2 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.343049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.343049Z digest=sha256:8d9dc9e10625c595f0217bd67850005207ca550b61de5801af6b7b07c3500322

Observation 5bd2b0d3-8bcc-4278-b687-13d9cf31877b · outbound

This paper cites Griffon v2: Advancing multimodal perception with high-resolution scaling and visual-language co-referring, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Griffon v2: Advancing multimodal perception with high-resolution scaling and visual-language co-referring, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.865686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.347776Z digest=sha256:851a41025fd217e461528fae1647c136831b2097096cf0f607e42ede86a87db7

Observation 05456998-e84d-446b-87f2-f5cab2bb312f · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Griffon: Spelling out all object locations at any granularity with large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.849590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.352754Z digest=sha256:7e18baa403052ed88a2fffd2559410f06063dcfc52d403c16f653ece67593308

Observation 5adb5b21-dda0-4a01-98fe-8b0fe5af8fa1 · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.357690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.357690Z digest=sha256:a65638c4509d3074a206338b45776edad397b2b0ba30a73122596a86baebe240

Observation 8ff683b7-901f-48ef-a96f-1e6668a92687 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.362335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.362335Z digest=sha256:634c79dd2a0cb70a330e1728337f8eded25213ea5d72e3061a80e0febc474be9

Observation acf63f30-37ca-40c5-8b35-65724dc9e9ef · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.367508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.367508Z digest=sha256:3b8435bbcfce79f01bfce22b93b930422ee693624f748341b67ec06e906bba01

Observation acaece5a-8065-43db-ae5e-13efe8aae67f · outbound

This paper cites Detrs beat yolos on real-time object detection, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Detrs beat yolos on real-time object detection, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.833175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.372330Z digest=sha256:65101f0537131e6e38996cba72c20810146622c6de99f2104ebec7a20d51d514

Observation 841aa40c-cb43-4908-b970-6244ff139a4d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.377155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.377155Z digest=sha256:3c295fa0284f68a02622df49356f1cffe4fd0cd58953a2a298af4fcec228e01b

Observation f1f79241-7ff3-4be4-a236-f2768dab37fe · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Deformable detr: Deformable transformers for end-to-end object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.816775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.381597Z digest=sha256:810f5b2e9f74dc719b839927d05c31f647e807fb46928115dc4c9f975785e124

Observation ae4fb308-6e4a-4e4b-92ce-a1acdf958aef · outbound

This paper cites Multi-step.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Multi-step

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:20:41.800601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.385870Z digest=sha256:5a64fcd418589c14594a0f31cb52659f20b9b1f9df61b58eb7cce571c6cb3214

Pith citing papers

No inbound Pith citation observations are available.