Pith. sign in

Paper Citation Record · LEDGER

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

As of 15 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2507.06272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06272 v3

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:03.038918Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.109470Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved40
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07667d61-3b90-4a50-90bc-9f7a8a79023a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.560688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.560688Z digest=sha256:b221b704c4164f6d29f37c62d6f801c9f18afeda93b4f89f544c92047e66490e

Observation c1f30626-190d-43bc-819e-4e08e8f582bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.566717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.566717Z digest=sha256:e3aa509d25edd371afce05c2b86bc91195d4f8e462bcd34a6aca961dd2f3fc76

Observation d23a0cfd-ddfe-4b24-8f4c-4404ada3607a · outbound

This paper cites Internlm2 technical report, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Internlm2 technical report, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.472428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.573028Z digest=sha256:0e94c17d696f30ec5d9cd1cbe0cdbeffab8f13bddb4a004c5cfc48b017246085

Observation 123fb9d3-7959-44c2-9874-ddae3b8b1325 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.578412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.578412Z digest=sha256:7f438b36e0acaee3b6e54329e1136ace401c5c5b99c234097c70d1a018189462

Observation c3620d91-5c04-46bf-893c-40517ccf4caa · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.585377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.585377Z digest=sha256:c6aa635ed8121a65e99fcd012f36f56c68d5ab40b24898520ed97d0c7d4f833f

Observation 946f5be3-94c9-4f85-bc6b-7c95eeeabba7 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.592591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.592591Z digest=sha256:b1bca518bf95ecb63c47116e6c5b0d4e18e03e5a151b0e5f9818f7432fd9f255

Observation d81d9a75-d6a1-44b0-939c-e8ea40a342be · outbound

This paper cites Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.454166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.600401Z digest=sha256:3654b6baaed598cb95f863c16061cf1573b3c940a42123aceac7370845295868

Observation b0dd2dbc-8952-4815-99c9-5af96197e5de · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.612441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.612441Z digest=sha256:ce81e75f71b2ad871a7d84a65b1c1e169b35b53e4d8936ce3c47ff25eb5be61d

Observation c8a9d86c-06d2-405c-bf13-9c82085b0b48 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.620079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.620079Z digest=sha256:7ce94bb3a7615df75b165ad9a83d32b77d39c15a5c61b417510b2d4316248e52

Observation 15eff456-804c-4064-8188-8081a0d0bf2d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.433594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.626331Z digest=sha256:7094c98b7086e026fb6104096ba85041c242190189f806a1aff2e612a0340353

Observation 1c9bccd3-d49d-4815-ac94-3b5dbda9a750 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Vizwiz grand challenge: Answering visual questions from blind people

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.632584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.632584Z digest=sha256:c946eef58866393247248854a8b5eeceb60c92247b08bb5d423ccfef6f03a748

Observation d81508aa-362c-4002-8d5a-19dc7109a3c5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cogagent: A visual language model for gui agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.402511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.639805Z digest=sha256:8afedd72488d22ae78de22d2783c366ee8b6f6e9a78782e8be2e11c538312ead

Observation 7d658121-544e-4f2b-8fa0-5497626ca68c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lora: Low-rank adaptation of large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.385021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.645404Z digest=sha256:e7c361539ab33cee85b9a7288e593b58ac6aabae480a4e4dcec7f7bbfc8e8d15

Observation 2fd3d24c-5388-4590-a076-c05969ca92b0 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.650218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.650218Z digest=sha256:3b197bdd4262bfc94922f6699128445fd66271d9e5751ef6209e1b57f393e16c

Observation 92808272-303a-4110-968f-bc3d929e9205 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.367323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.655642Z digest=sha256:7992d4b6aac7bfd30c5d876ca3bd07df4c92df2619f996eef3ab628353569b1d

Observation 5d63ed73-3ef4-4c6e-8065-d8fcd8893485 · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Dvqa: Understanding data visualizations via ques- tion answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.660377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.660377Z digest=sha256:078ccacd43f425822d34e23e27eafc3d7bf829731eb5ae4d8810b74373d52886

Observation be61820e-baf0-4138-88c9-9fe5c3de2be3 · outbound

This paper cites A diagram is worth a dozen images.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A diagram is worth a dozen images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.666783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.666783Z digest=sha256:bc5210fd78d8635efe58ef8888ff9b407f1a8578364f931515e8237de50053ee

Observation 5f422613-5648-437a-a5b1-21c108970889 · outbound

This paper cites Segment any- thing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Segment any- thing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.679727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.679727Z digest=sha256:d55fc8de1a96e1e5e6d58c1023801003cfcfa9c8b893bc8c9596a40d696da022

Observation f9151211-5ed7-40c3-a995-6a6ec82e37ca · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lisa: Reasoning segmentation via large language model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.296163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.686421Z digest=sha256:945a33a6c164af75af04c929a4fc9c7a20d13ffa240e224f06bb3b9b77b8f1b5

Observation 9912157e-0ba8-4078-98f6-3d9d0e1c2f97 · outbound

This paper cites Text4Seg: Reimagining Image Segmentation as Text Generation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Text4Seg: Reimagining Image Segmentation as Text Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.692538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.692538Z digest=sha256:2b419bd583f8a70ebfe0f405a3eadeb4124eab6ecb4cdbae884fd3f6a9cf9c78

Observation 438b3181-0746-4800-8651-c1a3abe12590 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.699518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.699518Z digest=sha256:11c689ddb38857f90e18466c8fd7fa83f85318d9a834a4c14459e98e68f97042

Observation e96330e5-cbf0-4c07-b9e9-957de476c346 · outbound

This paper cites OMG-Seg: Is One Model Good Enough For All Segmentation?.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance OMG-Seg: Is One Model Good Enough For All Segmentation?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.706354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.706354Z digest=sha256:4fea8709fa276293be90645b154b7af8cd3bb78fbfc3f5ef26cb15baafeaefe9

Observation 6c3b72e5-80cd-47c4-ade6-52e4b22c0a6a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Evaluating object hallucination in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.267371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.713875Z digest=sha256:3a7dd3942e943d792c2a1efb063806f867373c11f955f1a2e088cc2634880ae2

Observation c1effccf-7352-46f9-ac3a-7c10cf449f54 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.719308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.719308Z digest=sha256:7b7086267f4b4ab6dff3c4669812e7ea97cf34ab08497c893dafe84ec8f63ea2

Observation 776f9d73-49bf-4390-8d13-8be844bf18ac · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.248510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.725997Z digest=sha256:47b424d0c48eb29758f6b892c90e97e335c596a5dc8e77c8b1d97db93966fac3

Observation 66c9b9c1-9e14-4c63-98cd-3117bb196e78 · outbound

This paper cites Gres: Gener- alized referring expression segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gener- alized referring expression segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.231802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.731847Z digest=sha256:e8354dea6d5f1c1d49575f8a088ae10259c56d8aa850205ad5e3c413a09054ee

Observation 1309b116-be41-4492-8150-93af04f8a6ef · outbound

This paper cites Gres: Gen- eralized referring expression segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gen- eralized referring expression segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.214525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.737729Z digest=sha256:0b897c75cc00e6eaef9abcc8071a8d9aa43f9bfe3161643c415c407dc4bc572b

Observation 25cbb403-d745-40c7-9d46-33681c38d58f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.197835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.743293Z digest=sha256:0cecf5545c8fd2a9ff5b16edfe06294ecb82876d24d9fb8cb8b94919cd4de9a2

Observation 3a6f45e1-600a-4186-b0b4-b66d8c1959aa · outbound

This paper cites Visual instruction tuning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.181474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.750364Z digest=sha256:1372559b3f3360fc9af9be3eb4013889894c4993ca79bac44e57fb340eeb768e

Observation b6798a7d-cf83-4917-bc88-4894cf994b78 · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.756176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.756176Z digest=sha256:f81faaad5fc941349f239cb9195a324c21eff727fbcb8ee4205509a13e32c060

Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.762351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.762351Z digest=sha256:efa71cd0a616e1554da6ba0811b8fbcf0850bbb2c9cc4ebcc1ec6ab7e161ddc8

Observation 608ab5c5-c239-4760-b351-dfcd833a97c9 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.768004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.768004Z digest=sha256:3436c20f0b91d242319703787b4aba47b91d7785fb0a0f284549067b36d08e9c

Observation 4b577851-b868-4a87-a58d-e891afe63e52 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.773169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.773169Z digest=sha256:f1c03abff992fd41f54146a73b51d2457d76def2b0b4db5831837b102a9493f5

Observation 2e4a3485-1b56-4e99-860b-5402c90fa88f · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.779348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.779348Z digest=sha256:73148a76300991d24dc0a4aa9d9510f749c390b0f6da428c7bcb9122b5a23cb5

Observation be5cb8d7-a020-43c8-a85c-1bb846236cd4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.152584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.786553Z digest=sha256:6df6edb94fd75cc34bdc39b5710146c4cadc2156b24010fdd0968fb9b0ea2541

Observation c64fb980-aa3b-45df-be78-a21e7f0a2fe3 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.794185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.794185Z digest=sha256:de3e5f5ec6aef627529fc2f3d0a4b797e40e7c1d6181872c1ed27589b43c4901

Observation 327d3d31-edab-4095-808e-0eb026ee171d · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.135326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.800568Z digest=sha256:6f38ef24e213b0ab706e908c0d5e91e88db4b194541a1eda9c072332e69c9c2a

Observation 5625cbc4-79a1-42c8-8255-c8f792690acb · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.806358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.806358Z digest=sha256:1f2f3742e903d717a470abf74e7da15f31d64ae017e2f90724c1baa6c396eb80

Observation 446b98f5-a107-4b23-8d14-f98ce24d6a7b · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mod- eling context between objects for referring expression under- standing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.811525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.811525Z digest=sha256:f8a14ab5209cd788fa5a186767238abb16ce0fdec2bc7a9790d368c4888f7cbd

Observation a79209ef-18ee-4598-bffc-0ed393c22227 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.816755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.816755Z digest=sha256:6c7a03d863714a6d1c91de996a994060295fe9923d21b58802c182741a7d911f

Observation 087fa912-cbf8-437d-bf28-1d9225376729 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.822075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.822075Z digest=sha256:7ab4f98fc1318c2e7bce8daa533ddd575cc016ecb0b0353db15a42fe4d933339

Observation 3ce61a4c-e2d4-4cfb-9076-cd599d3f71df · outbound

This paper cites Reasoning to attend: Try to understand how¡ seg¿ token works.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Reasoning to attend: Try to understand how¡ seg¿ token works

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.108283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.828356Z digest=sha256:c3a9257fea5f8d4d902252d77e59cf797bfaaa0b5216d68a6e950c06aea6a8a5

Observation f2eca709-119d-497e-88a1-f6efe28401e5 · outbound

This paper cites 10 Glamm: Pixel grounding large multimodal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance 10 Glamm: Pixel grounding large multimodal model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.091454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.834105Z digest=sha256:493519e511a8e47101b1d81bdc59e98555d317bde5a07b1dfb7c6f06c8b6ab3d

Observation 321b42bd-a7ad-405b-bdc7-4d7fd7a93ac3 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Pixellm: Pixel reasoning with large multimodal model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.075155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.839396Z digest=sha256:c79a54370eb2faad25b029a9704246b3f47c8d81972511ab21f71f97de0a0da6

Observation ef42190d-5df3-4e58-8f53-d637060f1479 · outbound

This paper cites Object hallucination in image cap- tioning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Object hallucination in image cap- tioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.846541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.846541Z digest=sha256:37520bdb1b7e83c107d968e00a4720be99f5814b16cb49901d95e2b6ec7c3491

Observation 08f1824d-f5a8-4416-970e-15639a0b1d05 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.047364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.851753Z digest=sha256:e7ffaffe03279598e98d48038048e36c1f89ad64710c4486d9377eec91f4ddb7

Observation 70e555ba-6fe2-4a70-b233-858ccd208a78 · outbound

This paper cites Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.027874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.856963Z digest=sha256:0c03909ea4d68e2a53bff473553ce464f35d84ef537ec658fa5397003de96ec2

Observation a1553ffa-e220-4930-84d1-3e7668aa78e4 · outbound

This paper cites Towards vqa models that can read.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.862897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.862897Z digest=sha256:4477a2f76b60b3f9487fabc46f00e7f71573a8a6e365c57f28fbdf0356f6abf0

Observation 951a132c-acf1-4cbd-b158-19561d0a73d8 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.997513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.869346Z digest=sha256:95bff6aeed209cad7aaa2470f5e532afe105adb3783d92e7ff879df12f0e9da1

Observation 34dc3d2c-51ae-41a9-a577-67bf4810d1a9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.874944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.874944Z digest=sha256:6d24713975afc73f9df39d6dfdf257512b9e804e7c85b761adc3e84e0115ddf1

Observation 1834aff1-5fff-49d6-8c83-2fbb9966a293 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.881306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.881306Z digest=sha256:c73a31811acec43752dbb5f287761c8dd5d7c30b49cf696d079d5e6e3aa9e27b

Observation 6664a6e4-dea9-464e-bce3-8631bbe9691b · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visionllm: Large language model is also an open- ended decoder for vision-centric tasks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.980701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.892908Z digest=sha256:57ae3d20f3c837caf49370ac7e6e2354732065c7fc489004e2e5e10ec4fb736f

Observation 7727ef6d-8e2a-4681-add6-7b3405fa1997 · outbound

This paper cites Hierar- chical open-vocabulary universal image segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Hierar- chical open-vocabulary universal image segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.963856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.899849Z digest=sha256:c1ac65e94d115d9b69329c6aba34f2f4a94e07d03cdc8768735759fe04b69d22

Observation a118ef6e-a4e0-4a5f-a839-c70b4d9bbcac · outbound

This paper cites SegLLM: Multi-round Reasoning Segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance SegLLM: Multi-round Reasoning Segmentation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.905453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.905453Z digest=sha256:a90f2e7facafe3ea8300e0e73ba2bcdcf5cdc6c27ab6b23ee672f437d5112717

Observation 079b2c5b-371f-4f8f-8e52-c5cff871c6e6 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.914864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.914864Z digest=sha256:9fcf295cc7b3c851f9c0fde02c12443e4208d64dbc00a6d83e4d7777045e6b4d

Observation 0b659962-2df3-4eea-b750-fe0c9decdd8a · outbound

This paper cites General object foundation model for images and videos at scale.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance General object foundation model for images and videos at scale

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.946172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.920393Z digest=sha256:8d097d2e1a1728b63ce61a73fdf273167d94f1b943bf02c478ca73f8df223e8b

Observation a8036437-e011-48ec-9078-abfa965eb13e · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.926461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.926461Z digest=sha256:a14ec529657dca30216d1ed4ae188ac3e7fcc8800dba6e995e6b6b6e064f67f8

Observation 59987610-2c72-4322-9dd0-1d553565e551 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance V*: Guided visual search as a core mechanism in multimodal llms

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.929391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.934981Z digest=sha256:ddf331ba6a81006221f0ee3bacddbef17f4db5763d9e2405e4af78b9b9e7ac04

Observation 30aae116-914d-4b0e-8ae2-252453f7df71 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gsva: Generalized segmentation via multimodal large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.911800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.940703Z digest=sha256:199951f786191ce7d14cda4ff1885bf37a5fe7194b3c7f964b5271fd195b9891

Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.946036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.946036Z digest=sha256:f2a055b24a5410312f46bb6ce0585c6f9e3d0d8d10505ea19593876bbbaf2cfe

Observation 991fa207-564b-42eb-b34a-12de5323ffce · outbound

This paper cites Qwen2 technical report, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2 technical report, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.894309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.952327Z digest=sha256:89fec812ee48f69bfcd6eed851da33bea6970823e09eb8fe7c15725217c1600f

Observation ae5f731d-8a5b-4e55-8359-535d567b62d7 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.959690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.959690Z digest=sha256:ac02aa93be20fdc005a5170e9566b242b09050263166557ee4bdbdee57a924e2

Observation 6b3102a1-ab8f-4168-8805-3f6862cb943c · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.876247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.966226Z digest=sha256:70146b82ada5dc201ba545375523e672b48042fe4933e00a12b4282c259d8af1

Observation f156745b-ab03-44ca-becf-967ee861c55b · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.971505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.971505Z digest=sha256:f5602df9ebbd9c4e9258abc00f3eb76ae24e28c22946d15ca0b80dc6648fc7fb

Observation 2182922a-1b65-4d5c-8bf2-36da9a7155b1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.977524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.977524Z digest=sha256:6641c6944663f28b8235d666e2206e4923a798044b15ae9a84a60df8f6634bc0

Observation 836ac2fe-de2b-4583-aa04-dd43e5df3ee7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.857363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.984296Z digest=sha256:2315424f7bead41886a56b16bad221eaa8ada51dba33acafd9c381c1eeecada1

Observation fe05fdce-4e5e-4b39-bd41-193b88ab0b47 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ferret: Refer and ground anything anywhere at any granularity

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.839191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.991047Z digest=sha256:ba91ded41af103e4a9bc76a820d8184ada15eb2e1fabd1132661271daaac81b0

Observation 19788345-0a4a-453f-8243-d4300b29a1ec · outbound

This paper cites Modeling context in referring expres- sions.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Modeling context in referring expres- sions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.822751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.998247Z digest=sha256:a690b27adc678ec10905fe51919f4f198e70ac5d1c959d17cc8aa668c7f9c64d

Observation b9a6494e-4048-411a-9dcd-9946a68d8d5b · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.005676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.005676Z digest=sha256:0c63b61c39c7b1f5e65f0acbb53bd514296f14841ea1ca8652a59410fc89c1c6

Observation fa0bda23-d1d4-4f18-ab4c-7c1957fad63b · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.012476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.012476Z digest=sha256:e6a778cb88542030bcb952d3cceff4ea7b80add14ace5797fff8387c6cf59b61

Observation 5faef72c-b81c-4aea-8c07-04ecd20f71e1 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.805984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:03.020693Z digest=sha256:d387a12e62612d14e63668a222a01e28328c5fd6b5e2a5d2a6bfde6ff0923bb0

Observation d06d3946-b1e1-46c6-82e1-dcc182433d74 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Psalm: Pixelwise segmentation with large multi-modal model

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.787981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:03.026335Z digest=sha256:7cc16f83f7748aff272b2c2374feb5e95db4aa0dac9a42e60f0d1f79b253d062

Observation 068d6238-60ab-4185-b6b0-29f180f24a04 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.032130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.032130Z digest=sha256:c692201f30f2a9cd5674b0817732cad39559fb0cdc94ba3a1d82357ed2afc8d9

Observation ffabad69-66d6-401c-9b0b-cfdb05f446f3 · outbound

This paper cites When the provided information is insufficient, respond with ‘Unanswerable,’.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance When the provided information is insufficient, respond with ‘Unanswerable,’

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.769446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:03.038918Z digest=sha256:88d398a2cfd91b37d8d67c788640e070583e201171dbfcea7df0c768ad1a92b6

Observation fd18e0b8-357d-4092-a28a-536a65f6a9a0 · outbound

This paper cites an unresolved cited work.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T19:25:04.323807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:25:02.672598Z digest=sha256:d7f9e2e32e11ca3250480ed199eea630df568dbc89b0d7047b8c619e4e5b321e

Pith citing papers

Observation fa7ad16a-b5f6-4ed2-bd67-b6d1ce5d30c6 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.111055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:036736c578f809e54237dcfbf4513936e84ed642ad1b9f369eda06fc2c109753