Pith. sign in

Paper Citation Record · LEDGER

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

As of 15 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2507.06272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06272 v3

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:03.038918Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.109470Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved40
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07667d61-3b90-4a50-90bc-9f7a8a79023a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.560688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.560688Z digest=sha256:b221b704c4164f6d29f37c62d6f801c9f18afeda93b4f89f544c92047e66490e

Observation c1f30626-190d-43bc-819e-4e08e8f582bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.566717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.566717Z digest=sha256:e3aa509d25edd371afce05c2b86bc91195d4f8e462bcd34a6aca961dd2f3fc76

Observation d23a0cfd-ddfe-4b24-8f4c-4404ada3607a · outbound

This paper cites Internlm2 technical report, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Internlm2 technical report, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.472428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.573028Z digest=sha256:48e07aff7c5def7109478a7e2ee5436a0f86fd6cae9690bbc85d15dc55a698a3

Observation 123fb9d3-7959-44c2-9874-ddae3b8b1325 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.578412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.578412Z digest=sha256:7f438b36e0acaee3b6e54329e1136ace401c5c5b99c234097c70d1a018189462

Observation c3620d91-5c04-46bf-893c-40517ccf4caa · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.585377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.585377Z digest=sha256:c6aa635ed8121a65e99fcd012f36f56c68d5ab40b24898520ed97d0c7d4f833f

Observation 946f5be3-94c9-4f85-bc6b-7c95eeeabba7 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.592591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.592591Z digest=sha256:b1bca518bf95ecb63c47116e6c5b0d4e18e03e5a151b0e5f9818f7432fd9f255

Observation d81d9a75-d6a1-44b0-939c-e8ea40a342be · outbound

This paper cites Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.454166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.600401Z digest=sha256:af5be149718336dcc4e9f58d5428586037009e2528c97c566d5b0012c77c7e95

Observation b0dd2dbc-8952-4815-99c9-5af96197e5de · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.612441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.612441Z digest=sha256:ce81e75f71b2ad871a7d84a65b1c1e169b35b53e4d8936ce3c47ff25eb5be61d

Observation c8a9d86c-06d2-405c-bf13-9c82085b0b48 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.620079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.620079Z digest=sha256:7ce94bb3a7615df75b165ad9a83d32b77d39c15a5c61b417510b2d4316248e52

Observation 15eff456-804c-4064-8188-8081a0d0bf2d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.433594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.626331Z digest=sha256:c68fb0e6e1e5eaaaf172713f3fac24fb0d918a0340cfa3afd2911af6c80bacbb

Observation 1c9bccd3-d49d-4815-ac94-3b5dbda9a750 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Vizwiz grand challenge: Answering visual questions from blind people

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.632584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.632584Z digest=sha256:c946eef58866393247248854a8b5eeceb60c92247b08bb5d423ccfef6f03a748

Observation d81508aa-362c-4002-8d5a-19dc7109a3c5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cogagent: A visual language model for gui agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.402511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.639805Z digest=sha256:d45c7d640aa2d7eafbab2f9f051e1b8e2f987efa9c86612658e0918da58fb733

Observation 7d658121-544e-4f2b-8fa0-5497626ca68c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lora: Low-rank adaptation of large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.385021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.645404Z digest=sha256:e68354097f93232ef6ac44cf3e913884b83eb91075f04164362b88cf04b0a67b

Observation 2fd3d24c-5388-4590-a076-c05969ca92b0 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.650218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.650218Z digest=sha256:3b197bdd4262bfc94922f6699128445fd66271d9e5751ef6209e1b57f393e16c

Observation 92808272-303a-4110-968f-bc3d929e9205 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.367323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.655642Z digest=sha256:67d3e2cdeb6ddb8f1008ac91894b9ef27bb60fd074463c97f37ceb6011c56562

Observation 5d63ed73-3ef4-4c6e-8065-d8fcd8893485 · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Dvqa: Understanding data visualizations via ques- tion answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.660377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.660377Z digest=sha256:078ccacd43f425822d34e23e27eafc3d7bf829731eb5ae4d8810b74373d52886

Observation be61820e-baf0-4138-88c9-9fe5c3de2be3 · outbound

This paper cites A diagram is worth a dozen images.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A diagram is worth a dozen images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.666783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.666783Z digest=sha256:bc5210fd78d8635efe58ef8888ff9b407f1a8578364f931515e8237de50053ee

Observation 5f422613-5648-437a-a5b1-21c108970889 · outbound

This paper cites Segment any- thing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Segment any- thing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.679727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.679727Z digest=sha256:d55fc8de1a96e1e5e6d58c1023801003cfcfa9c8b893bc8c9596a40d696da022

Observation f9151211-5ed7-40c3-a995-6a6ec82e37ca · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lisa: Reasoning segmentation via large language model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.296163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.686421Z digest=sha256:69ebc50f1dc52dcebb57f3d4abc4e34b554f527b2932697c41380db9ea6e71cf

Observation 9912157e-0ba8-4078-98f6-3d9d0e1c2f97 · outbound

This paper cites Text4Seg: Reimagining Image Segmentation as Text Generation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Text4Seg: Reimagining Image Segmentation as Text Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.692538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.692538Z digest=sha256:2b419bd583f8a70ebfe0f405a3eadeb4124eab6ecb4cdbae884fd3f6a9cf9c78

Observation 438b3181-0746-4800-8651-c1a3abe12590 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.699518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.699518Z digest=sha256:11c689ddb38857f90e18466c8fd7fa83f85318d9a834a4c14459e98e68f97042

Observation e96330e5-cbf0-4c07-b9e9-957de476c346 · outbound

This paper cites OMG-Seg: Is One Model Good Enough For All Segmentation?.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance OMG-Seg: Is One Model Good Enough For All Segmentation?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.706354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.706354Z digest=sha256:4fea8709fa276293be90645b154b7af8cd3bb78fbfc3f5ef26cb15baafeaefe9

Observation 6c3b72e5-80cd-47c4-ade6-52e4b22c0a6a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Evaluating object hallucination in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.267371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.713875Z digest=sha256:729c7334ae6c48da754179053b68d15cfd7227b96e1df7c199c343249aec54fb

Observation c1effccf-7352-46f9-ac3a-7c10cf449f54 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.719308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.719308Z digest=sha256:7b7086267f4b4ab6dff3c4669812e7ea97cf34ab08497c893dafe84ec8f63ea2

Observation 776f9d73-49bf-4390-8d13-8be844bf18ac · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.248510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.725997Z digest=sha256:68e5adacac84fcf0c4ea29cca32fff0a5dda4040bb7b7b5b5d1a93ec994d1e78

Observation 66c9b9c1-9e14-4c63-98cd-3117bb196e78 · outbound

This paper cites Gres: Gener- alized referring expression segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gener- alized referring expression segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.231802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.731847Z digest=sha256:e072aebe05f0eef6dcd197e7e5ab75a507190f70eb6a0b33a5606b511b6ff7bc

Observation 1309b116-be41-4492-8150-93af04f8a6ef · outbound

This paper cites Gres: Gen- eralized referring expression segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gen- eralized referring expression segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.214525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.737729Z digest=sha256:0f9468e91ba4e139a62037dd5addb0e4c5ab744ed67e10bd136d88c2e79ac072

Observation 25cbb403-d745-40c7-9d46-33681c38d58f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.197835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.743293Z digest=sha256:bcf8db7a87f6352413499c9693d8906b0f6f44c7e26656261af1b2c0c373e615

Observation 3a6f45e1-600a-4186-b0b4-b66d8c1959aa · outbound

This paper cites Visual instruction tuning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.181474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.750364Z digest=sha256:acaa9847abf30adfcfdb8ff0c8a28d04b9336720308f2b68d7aa4dd9c46d072c

Observation b6798a7d-cf83-4917-bc88-4894cf994b78 · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.756176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.756176Z digest=sha256:f81faaad5fc941349f239cb9195a324c21eff727fbcb8ee4205509a13e32c060

Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.762351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.762351Z digest=sha256:efa71cd0a616e1554da6ba0811b8fbcf0850bbb2c9cc4ebcc1ec6ab7e161ddc8

Observation 608ab5c5-c239-4760-b351-dfcd833a97c9 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.768004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.768004Z digest=sha256:3436c20f0b91d242319703787b4aba47b91d7785fb0a0f284549067b36d08e9c

Observation 4b577851-b868-4a87-a58d-e891afe63e52 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.773169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.773169Z digest=sha256:f1c03abff992fd41f54146a73b51d2457d76def2b0b4db5831837b102a9493f5

Observation 2e4a3485-1b56-4e99-860b-5402c90fa88f · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.779348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.779348Z digest=sha256:73148a76300991d24dc0a4aa9d9510f749c390b0f6da428c7bcb9122b5a23cb5

Observation be5cb8d7-a020-43c8-a85c-1bb846236cd4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.152584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.786553Z digest=sha256:46133ff6a9d51d5ae82cbaf25031d6174d88552be4eaf166e2005fa8a34426e2

Observation c64fb980-aa3b-45df-be78-a21e7f0a2fe3 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.794185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.794185Z digest=sha256:b0e1a05a283a43fc41e380627b59f2e899e7360f6cd29888797c1268fddac1c4

Observation 327d3d31-edab-4095-808e-0eb026ee171d · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.135326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.800568Z digest=sha256:9928e1776a40a5ea416c1484e91ff67d6b062ec6e802a47a878cf3a03df58878

Observation 5625cbc4-79a1-42c8-8255-c8f792690acb · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.806358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.806358Z digest=sha256:1f2f3742e903d717a470abf74e7da15f31d64ae017e2f90724c1baa6c396eb80

Observation 446b98f5-a107-4b23-8d14-f98ce24d6a7b · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mod- eling context between objects for referring expression under- standing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.811525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.811525Z digest=sha256:f8a14ab5209cd788fa5a186767238abb16ce0fdec2bc7a9790d368c4888f7cbd

Observation a79209ef-18ee-4598-bffc-0ed393c22227 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.816755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.816755Z digest=sha256:6c7a03d863714a6d1c91de996a994060295fe9923d21b58802c182741a7d911f

Observation 087fa912-cbf8-437d-bf28-1d9225376729 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.822075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.822075Z digest=sha256:7ab4f98fc1318c2e7bce8daa533ddd575cc016ecb0b0353db15a42fe4d933339

Observation 3ce61a4c-e2d4-4cfb-9076-cd599d3f71df · outbound

This paper cites Reasoning to attend: Try to understand how¡ seg¿ token works.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Reasoning to attend: Try to understand how¡ seg¿ token works

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.108283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.828356Z digest=sha256:0466ac86919267497d675daf6d6333a8f5e4f73d677e8193dd1ae3ca6dd813e3

Observation f2eca709-119d-497e-88a1-f6efe28401e5 · outbound

This paper cites 10 Glamm: Pixel grounding large multimodal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance 10 Glamm: Pixel grounding large multimodal model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.091454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.834105Z digest=sha256:7bd0cb2484aa1e271315a6baa2fb38adb672900e783d83cefe5d5b08054e597a

Observation 321b42bd-a7ad-405b-bdc7-4d7fd7a93ac3 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Pixellm: Pixel reasoning with large multimodal model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.075155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.839396Z digest=sha256:c07e68d3ad76f2255b1559397bc892ee4515b331bc3f25e0603a603efe4dd087

Observation ef42190d-5df3-4e58-8f53-d637060f1479 · outbound

This paper cites Object hallucination in image cap- tioning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Object hallucination in image cap- tioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.846541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.846541Z digest=sha256:37520bdb1b7e83c107d968e00a4720be99f5814b16cb49901d95e2b6ec7c3491

Observation 08f1824d-f5a8-4416-970e-15639a0b1d05 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.047364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.851753Z digest=sha256:85818036b36ca6d5c544ad8d0a1f46f27b94ad715e7b1bab5cb3e2ba45b382c9

Observation 70e555ba-6fe2-4a70-b233-858ccd208a78 · outbound

This paper cites Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.027874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.856963Z digest=sha256:2c0a073fe08f797e9131b5bc63b48dc0f7e9d187c901d7ec601c7929d5809d32

Observation a1553ffa-e220-4930-84d1-3e7668aa78e4 · outbound

This paper cites Towards vqa models that can read.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.862897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.862897Z digest=sha256:4477a2f76b60b3f9487fabc46f00e7f71573a8a6e365c57f28fbdf0356f6abf0

Observation 951a132c-acf1-4cbd-b158-19561d0a73d8 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.997513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.869346Z digest=sha256:0beeb50d787d32aaac3060e8f328baf63320ed9df984876eb8a0995692e5bba2

Observation 34dc3d2c-51ae-41a9-a577-67bf4810d1a9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.874944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.874944Z digest=sha256:6d24713975afc73f9df39d6dfdf257512b9e804e7c85b761adc3e84e0115ddf1

Observation 1834aff1-5fff-49d6-8c83-2fbb9966a293 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.881306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.881306Z digest=sha256:c73a31811acec43752dbb5f287761c8dd5d7c30b49cf696d079d5e6e3aa9e27b

Observation 6664a6e4-dea9-464e-bce3-8631bbe9691b · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visionllm: Large language model is also an open- ended decoder for vision-centric tasks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.980701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.892908Z digest=sha256:8687ac5858b1dd530a8b82776e5f0b082dd2f7476bf30d6109e5b9fee1c44de9

Observation 7727ef6d-8e2a-4681-add6-7b3405fa1997 · outbound

This paper cites Hierar- chical open-vocabulary universal image segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Hierar- chical open-vocabulary universal image segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.963856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.899849Z digest=sha256:9b3aef1d44817504d54074ce82eea533f6aa2d9f8cf3ea914d42c159f33f2e6d

Observation a118ef6e-a4e0-4a5f-a839-c70b4d9bbcac · outbound

This paper cites SegLLM: Multi-round Reasoning Segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance SegLLM: Multi-round Reasoning Segmentation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.905453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.905453Z digest=sha256:a90f2e7facafe3ea8300e0e73ba2bcdcf5cdc6c27ab6b23ee672f437d5112717

Observation 079b2c5b-371f-4f8f-8e52-c5cff871c6e6 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.914864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.914864Z digest=sha256:9fcf295cc7b3c851f9c0fde02c12443e4208d64dbc00a6d83e4d7777045e6b4d

Observation 0b659962-2df3-4eea-b750-fe0c9decdd8a · outbound

This paper cites General object foundation model for images and videos at scale.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance General object foundation model for images and videos at scale

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.946172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.920393Z digest=sha256:40b5f187b0b5925b1b1c384e6bb38d11a3e99f1599483f7a6bccc379764ee673

Observation a8036437-e011-48ec-9078-abfa965eb13e · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.926461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.926461Z digest=sha256:a14ec529657dca30216d1ed4ae188ac3e7fcc8800dba6e995e6b6b6e064f67f8

Observation 59987610-2c72-4322-9dd0-1d553565e551 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance V*: Guided visual search as a core mechanism in multimodal llms

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.929391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.934981Z digest=sha256:53c05de98c5982afd7ef1a08b81caaeff1abb159add8507ec0067bb8c850db8e

Observation 30aae116-914d-4b0e-8ae2-252453f7df71 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gsva: Generalized segmentation via multimodal large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.911800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.940703Z digest=sha256:564ad8b316b19b281f921e6f40e00a6297846eb32790a0c8d2a46ff3aef63e37

Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.946036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.946036Z digest=sha256:f2a055b24a5410312f46bb6ce0585c6f9e3d0d8d10505ea19593876bbbaf2cfe

Observation 991fa207-564b-42eb-b34a-12de5323ffce · outbound

This paper cites Qwen2 technical report, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2 technical report, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.894309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.952327Z digest=sha256:9de2e32a33d64730099d4e2c68f95dbcc4123c7abf95097e58f9e1171628e9ad

Observation ae5f731d-8a5b-4e55-8359-535d567b62d7 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.959690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.959690Z digest=sha256:ac02aa93be20fdc005a5170e9566b242b09050263166557ee4bdbdee57a924e2

Observation 6b3102a1-ab8f-4168-8805-3f6862cb943c · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.876247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.966226Z digest=sha256:7f27c58f4cc1d88459af3cc0ada84dbdd440e66ea5b18ed6e76d5a9f115c4799

Observation f156745b-ab03-44ca-becf-967ee861c55b · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.971505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.971505Z digest=sha256:f5602df9ebbd9c4e9258abc00f3eb76ae24e28c22946d15ca0b80dc6648fc7fb

Observation 2182922a-1b65-4d5c-8bf2-36da9a7155b1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.977524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.977524Z digest=sha256:6641c6944663f28b8235d666e2206e4923a798044b15ae9a84a60df8f6634bc0

Observation 836ac2fe-de2b-4583-aa04-dd43e5df3ee7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.857363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.984296Z digest=sha256:2fa226053026b394093d463104fe5ba88e5cc9b3b20dc76dafefb8ca36ce5ce7

Observation fe05fdce-4e5e-4b39-bd41-193b88ab0b47 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ferret: Refer and ground anything anywhere at any granularity

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.839191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.991047Z digest=sha256:c468248145a04f95a7b39b58c51461cfb9f81a5bd24871806f224a22a31fb861

Observation 19788345-0a4a-453f-8243-d4300b29a1ec · outbound

This paper cites Modeling context in referring expres- sions.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Modeling context in referring expres- sions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.822751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.998247Z digest=sha256:ca9a00ade512460d47b2fc2d3641ef2b01b94f568d29044a84f87ed9d28f95af

Observation b9a6494e-4048-411a-9dcd-9946a68d8d5b · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.005676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.005676Z digest=sha256:0c63b61c39c7b1f5e65f0acbb53bd514296f14841ea1ca8652a59410fc89c1c6

Observation fa0bda23-d1d4-4f18-ab4c-7c1957fad63b · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.012476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.012476Z digest=sha256:e6a778cb88542030bcb952d3cceff4ea7b80add14ace5797fff8387c6cf59b61

Observation 5faef72c-b81c-4aea-8c07-04ecd20f71e1 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.805984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:03.020693Z digest=sha256:99139b26d2bbe731ddca0426733ed86c1bd1815debe9fa71d91e7d920ec66fab

Observation d06d3946-b1e1-46c6-82e1-dcc182433d74 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Psalm: Pixelwise segmentation with large multi-modal model

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.787981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:03.026335Z digest=sha256:8f84c77d892d20351437fb6c050e0616da5970d33ae6f8663e97b93fe0b45bc2

Observation 068d6238-60ab-4185-b6b0-29f180f24a04 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.032130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.032130Z digest=sha256:c692201f30f2a9cd5674b0817732cad39559fb0cdc94ba3a1d82357ed2afc8d9

Observation ffabad69-66d6-401c-9b0b-cfdb05f446f3 · outbound

This paper cites When the provided information is insufficient, respond with ‘Unanswerable,’.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance When the provided information is insufficient, respond with ‘Unanswerable,’

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.769446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:03.038918Z digest=sha256:cded84810406100774b4c3a943589278836cd3873570a389c29df265d3484528

Observation fd18e0b8-357d-4092-a28a-536a65f6a9a0 · outbound

This paper cites an unresolved cited work.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T19:25:04.323807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:25:02.672598Z digest=sha256:2c49148f994bc6fa8d748ea67df0c883f7b9a6e6ba001fe70d2add0fcfb3ae6b

Pith citing papers

Observation fa7ad16a-b5f6-4ed2-bd67-b6d1ce5d30c6 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.111055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:2747dcda1f7918080ab0a88607b1be81a99b5913a1a4b12251b083e604607973