Pith. sign in

Paper Citation Record · LEDGER

LlamaSeg: Image Segmentation via Autoregressive Mask Generation

As of 10 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2505.19422.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19422 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:58.783069Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T23:28:46.879238Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 623c40c9-214e-4fa2-a520-3b27fb5e0b64 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.765134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.765134Z digest=sha256:f2027136cd80e2bf1e0fedbe73f3449526b14827d8e5838b6e397e295d88bc70

Observation 9dcf6b2f-59a7-42ed-a7cb-9a1c7ec7e0f2 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:05.575503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:48.914096Z digest=sha256:43318644afbc73262e9901b1daa19f26c9d37474f591bd6542cb2dca005dab91

Observation ad4dcc7b-8d3c-4721-868a-1ac35484542d · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Emu: Generative Pretraining in Multimodality

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.060461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.060461Z digest=sha256:45c52e3833156a2b36bac7e753b9156b2609e047de728d4de3baf35450d9a95b

Observation 3953b53e-9e4c-4a3f-9a3f-b5b254c81aef · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visionllm: Large language model is also an open-ended decoder for vision-centric tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.248645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.248645Z digest=sha256:3a6c2caceec9c30eb5dd95cd90731562a16b65b1e54b03c498c5c0b0bce7dc2e

Observation 69a2b8a6-1ccb-4816-934a-7046e61ce2e3 · outbound

This paper cites u-llava: Unifying multi-modal tasks via large language model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation u-llava: Unifying multi-modal tasks via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:05.285437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:49.357629Z digest=sha256:e06f79e2029bcecf43231461b665b70a07b670595622e09cf6ae5ac077fc4eb6

Observation f390e4d9-901a-4de4-a06d-b1b5e7ec3a95 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:05.030481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:49.508094Z digest=sha256:57c2cb3f25750c6569169d7dbb0db667e2524867d3e07a73a036b6b2df795635

Observation 8704cb23-df55-4a38-8c1d-25b18d87cb06 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.688885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.688885Z digest=sha256:5df475a7979acc10991505f88cd56a6ac7515134e1772a8a96bb23dcbbe3c02d

Observation 32bea79f-30e9-455d-840f-11c34eafb280 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Lisa: Reasoning segmentation via large language model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.820249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.820249Z digest=sha256:67b35ff36623685839875d0f8371b768442abc55af32795f045d43fdd191e854

Observation 2c526dc8-57f1-4047-9239-89c4817eb0ef · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Glamm: Pixel grounding large multimodal model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.950170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.950170Z digest=sha256:930dd426be5fae453db03c15a694bc36b8e2c42f39be9e5a411cd61881895e3d

Observation 5430b6d1-ba48-4a40-8fdb-096520526187 · outbound

This paper cites Text4Seg: Reimagining Image Segmentation as Text Generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Text4Seg: Reimagining Image Segmentation as Text Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:50.127833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:50.127833Z digest=sha256:bfe6c894cd08a3036fa8eb0a6784e9a6109325cd5ec51ebda20cf0a150eb34de

Observation 76f4e1fc-a9d0-4370-b774-b79e204a20a1 · outbound

This paper cites Git: Towards generalist vision transformer through universal language interface.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Git: Towards generalist vision transformer through universal language interface

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:04.723255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:50.295316Z digest=sha256:cb3044e7245016d72543c44adaecd9291f86b9b816f38837dc3cadb9a6901d34

Observation 23f8c6b0-2607-469a-abe7-e9a42461feef · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Polyformer: Referring image segmentation as sequential polygon generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:04.457501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:50.454556Z digest=sha256:6bebc6cd6da576e719ab98ec42bb940ae1639b9dfe941ca423e7c1bd54ba53f7

Observation ce8ab02f-10ec-4b1e-97c8-8968d8748a48 · outbound

This paper cites Jack of all tasks master of many: Design- ing general-purpose coarse-to-fine vision-language model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Jack of all tasks master of many: Design- ing general-purpose coarse-to-fine vision-language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:04.159697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:50.622727Z digest=sha256:9a1796478090d173346e5e5f0dbd9ce0204d051cb9ad6bca36379f6ab0d59d95

Observation 6f1e769f-81df-49cc-945f-c5d91ad441a1 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Gsva: Generalized segmentation via multimodal large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.877209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:50.777068Z digest=sha256:b403f67c88f3de0a3d2269bbef2875aa97af04b360b93605890d23956b11a3c9

Observation a7b7b965-8d84-49f2-b43e-ed3c4c9d2b08 · outbound

This paper cites Segment anything.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Segment anything

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:50.924925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:50.924925Z digest=sha256:29591d127184fd7acb545170d6964d9c15094c7830993c5290bca09ecbbfc6ce

Observation a694bd22-80c1-4772-98f6-cb27f8dc1ff1 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Taming transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.038586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.038586Z digest=sha256:1ea23c3f48f7c50dd6566a2ce7f70a15076ecefec629a575aa2b116144501de6

Observation 38cbdd2d-192c-4889-b6a8-e42843b6f85b · outbound

This paper cites Zero-shot text-to-image generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Zero-shot text-to-image generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.218501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.218501Z digest=sha256:c338371bfd8353d49dedc70da50de3cc157040dac7b64f34182b94c92ab4e2a1

Observation 3aa89c91-a4c1-4e9d-924c-79ec84744fa5 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Generating diverse high-fidelity images with vq-vae-2

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.612443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:51.335353Z digest=sha256:da7cc21fe22a46f51133d1f72584d3c9423061f883d81aa7fc315946647d5af3

Observation 2754b107-31e3-4cd9-9572-3de20abc23c6 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.457069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.457069Z digest=sha256:4e56d96c5135e189ff6a9b8d0e5626b8eac7caefed0fa9481ac1d91bf0e2d920

Observation df5cb931-d91c-41d7-9709-d37dccc99a3b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.591027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.591027Z digest=sha256:ad406be56a0caf51d1dbda5a71c785bfc129e020f0a9e211f264c910d175e679

Observation 7bc3d74e-3129-4502-90d2-5e4a7c94b055 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.761106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.761106Z digest=sha256:0056454767ce6741c00ee8c6fae9df967b6b71ba0ba7fd8309ec3d0797f9f90a

Observation 7f91fb34-e4f6-4900-9543-f5120dcab32c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.877240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.877240Z digest=sha256:4988870d6eb9091a61a203917a186089225f31ae03bdd6ae3ac395fb0db9b4c2

Observation a10450b4-5556-455a-86d8-ef18a1e8fa32 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Learning transferable visual models from natural language supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.064538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.064538Z digest=sha256:472d6975fbf05b0274ee98cfb7b78d708a29aa18bdc2441a4d33d61c1a1267cb

Observation db3d08a4-bb40-4a73-b981-1a535471611c · outbound

This paper cites Attention is all you need.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Attention is all you need

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.224782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.224782Z digest=sha256:dbad9b0af7c049fb2d26e6d36182ac857206dd6b6452d671c678927723ed74f1

Observation b3c316b4-94f4-47dc-88d4-9ff1e9daa74b · outbound

This paper cites Improving language understanding by generative pre-training.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Improving language understanding by generative pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.336596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.336596Z digest=sha256:4e9c507fcb51ca1fa29f9008519f0cdaabfce06f4804a3f18321f673c3953ea1

Observation c4218a2b-b437-4ea8-90f8-21432545ca16 · outbound

This paper cites Language models are few-shot learners.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Language models are few-shot learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.477997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.477997Z digest=sha256:0639050990c5cdfabda82d408a9b686a0eb25cfcfe77c17ecfc9c166c7306ed6

Observation 28da60b8-52af-4cbb-aab5-c0c1a7597ab2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.589135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.589135Z digest=sha256:781665e67ed5b1ce16c3c3b27e254c953b253657839527eea41b3e3500af5d0d

Observation 7d3435b4-2434-4c01-b521-67d0c81911a6 · outbound

This paper cites Palm: Scaling language modeling with pathways.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Palm: Scaling language modeling with pathways

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.742201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.742201Z digest=sha256:07d3b63eb418414acfdbc6f4598c362c03ee6936e1f8392c35e9222dc182facd

Observation 4ffcabb1-f4f2-4be8-a4d5-82edf36b0c4f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.902659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.902659Z digest=sha256:f77f9c366255954f05fe5ebf74da114d829d7d065fa9de456278bd95a88af9eb

Observation 37bff697-7b0e-4354-a191-617cc4e3d316 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems , 36:34892–34916, 2023.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visual instruction tuning.Advances in neural information processing systems , 36:34892–34916, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.079223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.079223Z digest=sha256:9393450d13b6dc09089bc41577eb93d71075134fcabcc61c99d56dd675c306ba

Observation f3fc05cd-94a9-4966-a4be-fef7035f58d6 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.272244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.272244Z digest=sha256:f9a087af810659d90173ff2d22b31f5fe45ff947f0a16c6c7944cdf50f22bdc5

Observation 9c0e5efd-d244-4cbf-834a-2e22166b6fe6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.373312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.373312Z digest=sha256:de03c2f51118cca5d32ea357b8167542c61a535ea0817959b94de1c84259d59f

Observation 778a211f-b458-490b-8669-8af288f77ece · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.498373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.498373Z digest=sha256:2c788d17753b937761043bd075b808cc0ffc0a72adf10c522c2ea14cb0b97562

Observation 67503526-5313-413c-9f3b-e9846bdf4c35 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.654071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.654071Z digest=sha256:d989b1ef48dcb36dcf8489670577959ff7f1349684c97d1bb14eb2e12d247b40

Observation bc4dbb58-1eb3-4735-a636-3454e1bc0e3c · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Per-pixel classification is not all you need for semantic segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.335852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:53.841047Z digest=sha256:b64e89d0e0a7dafefd0708bddb2c3a4bf307352dc3a6b76796ad39fb3ac11fd9

Observation 19ede363-d424-4c61-add1-4794e31cb037 · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Segformer: Simple and efficient design for semantic segmentation with transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.067538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:53.964977Z digest=sha256:95e4707158dcb2054a48622f6ca57ad4663be04bd3f41466ab62d8b55f7780bc

Observation 86e01068-3a53-4f58-9c6f-1d3e6d932ddd · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:02.806222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:54.105161Z digest=sha256:08c0932c8f7ab49ef6c15961e594796b7d0c92c26eacbc5515f751612479a802

Observation 7040646a-69a7-4de8-94d8-54fc8a91c106 · outbound

This paper cites Vision-language transformer and query generation for referring segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Vision-language transformer and query generation for referring segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:02.504780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:54.320451Z digest=sha256:7191ada2d1dbce7114d768ea7f0b849d09903f2c721a5a4a5f0066ddf8de0a77

Observation 336a06b2-7b56-4ea6-844f-9e747b18ab12 · outbound

This paper cites Cris: Clip-driven referring image segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Cris: Clip-driven referring image segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:02.165503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:54.477855Z digest=sha256:258d1a613847c8a843fcb9d63186b7af396703ebe70a30e861d1218f77ea9eeb

Observation e8044630-6258-44fc-94f5-c111eaddc3cc · outbound

This paper cites Restr: Convolution- free referring image segmentation using transformers.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Restr: Convolution- free referring image segmentation using transformers

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:01.859025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:54.628990Z digest=sha256:5794fc146408c11025f44f8afd4ccbea5cb4df946aeb1cb79dedfd86401fa573

Observation a41cf8c1-127a-4aa8-bca1-6ada93eab595 · outbound

This paper cites Gres: Generalized referring expression segmen- tation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Gres: Generalized referring expression segmen- tation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:01.579004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:54.809115Z digest=sha256:8f96f4fb88d857f801dd91e3d47e9230953576da2c65f33142c565e7e8e00d60

Observation abd4dfea-31d3-4e2e-8979-37aecf7ebbfc · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Pixellm: Pixel reasoning with large multimodal model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:54.960384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:54.960384Z digest=sha256:63db6cfa0ec153b3f45c3b63e3fda823479486a46a631d1203a98fc7473150c9

Observation f127ebe1-442f-4a9c-8b4f-6c0ce9a41818 · outbound

This paper cites Generative semantic segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Generative semantic segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:01.224061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:55.096922Z digest=sha256:e0e3e38d922e074a38b919c33cc40758814bc74eedf35a4986c94256c56308fd

Observation c632f009-f28c-4ec8-a450-770c44e795a9 · outbound

This paper cites All in tokens: Unifying output space of visual tasks via soft token.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation All in tokens: Unifying output space of visual tasks via soft token

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:00.848461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:55.239331Z digest=sha256:870ecb4726b9ccd11315a24e837c7ea1af1d5df89e7b59be898fcc5a173d6da1

Observation e685d759-ba8f-47d8-b3a2-b88bacd1eddc · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.397888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.397888Z digest=sha256:7c0ded33c9df7064e029513fc5ca52269a1b2bcc7415b45b3801f5f74562b9a1

Observation 576d1f19-6221-437a-8a5b-eeba8c509471 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:00.528506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:55.554486Z digest=sha256:0330fd8ee2edeb490d5aca7c32ac393663d0f7ab51d319dac69a6916d6de6c65

Observation a7d720ca-b63c-439a-8d53-88cc3e268380 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.721329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.721329Z digest=sha256:0c2cc628b6a97d01c18b2d41680f4cb33a955141e7e16c9cfb73054b2156a22e

Observation 3ce10ca2-5e5e-45ba-a977-153754d25f46 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.907967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.907967Z digest=sha256:d7dd5b2a424c5b337c0ca33b97455c9584af253b7b94683bfa6f5fc10a1004a7

Observation c7b92e2a-86c0-4a27-aff9-bb5b3b3c847b · outbound

This paper cites Neural discrete representation learning.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Neural discrete representation learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.052079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.052079Z digest=sha256:301c738941493aefcbe6442f1ac33fe57bf0b3eb6a02bf43d4916e521e61dfce

Observation 4d710822-ad3f-410b-b1ee-8813fcc589ec · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Vector-quantized Image Modeling with Improved VQGAN

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.238340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.238340Z digest=sha256:e7ad736cc7ffdeec64152b408f9198438dc80b2bb09790e1e202370353bf3db6

Observation 2d1b0731-6331-444c-bf65-888d1f32a6d2 · outbound

This paper cites Movq: Modulating quantized vectors for high-fidelity image generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Movq: Modulating quantized vectors for high-fidelity image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:00.276993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:56.342837Z digest=sha256:f0e02297fe87f763cb1d28307be43d2a820577ab3d1f8b349c8e0b813b603513

Observation 1ad97a92-c8cf-40fb-97b2-e858d76d3741 · outbound

This paper cites Auto-encoding variational bayes, 2013.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Auto-encoding variational bayes, 2013

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.481060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.481060Z digest=sha256:f75643dad37e8b1900428a7542b072fde09fd9334bc242aa8e456c6c0e826ead

Observation 2fbe24c8-87c0-4efc-84f2-4d69eb3d4cce · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.599570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.599570Z digest=sha256:d1bd430206384cd6218136eb5f790086799077b7db3a086ad4e34d75fae4c92c

Observation 7f7fb99c-e21f-425a-8010-925854e1df01 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Gaussian Error Linear Units (GELUs)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.761381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.761381Z digest=sha256:5b01297e4c03c718ec207b2415f1a9db36746b90fe0dcd5b32abeabcaff46e2b

Observation 42735c2f-3bc0-46f5-a769-d5050f6abdcc · outbound

This paper cites Root mean square layer normalization.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Root mean square layer normalization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.940918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.940918Z digest=sha256:f6e4f5deeda218f44f6041455398a3540b30ddc26c8195a0a99292ec8b9ba640

Observation e693b3d5-37b3-47b8-8e2c-1867145ed260 · outbound

This paper cites GLU Variants Improve Transformer.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation GLU Variants Improve Transformer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.057555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.057555Z digest=sha256:603dbb02198e289018f01894cadc844e1aeff1de0a449a130eaad0645954da2d

Observation 406fd3ad-69bb-4710-a8f4-8fdf63e9dfd7 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Roformer: Enhanced transformer with rotary position embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.179181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.179181Z digest=sha256:c3288b5c0edb97198de5a2b4a4decfa0889ef71063a9fad6464f8884159cf0e0

Observation 3e7691ad-1c85-4d6a-8646-b65250830ad9 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.313128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.313128Z digest=sha256:ce419ca2b0831c9a8e31b052bba1c9e9fbefe43ac17035436bc608714e364480

Observation cef5d445-2c91-408e-ae4f-e378343b352d · outbound

This paper cites Sigmoid loss for language image pre-training.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Sigmoid loss for language image pre-training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.445520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.445520Z digest=sha256:cfcc69c218bbe45c2453fef97fd5d1be0facc574d17cdd5c4c2841785e937257

Observation 1a297325-00bb-422e-8fae-146270f85ee3 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Coco-stuff: Thing and stuff classes in context

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:59.900587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:57.626054Z digest=sha256:408426153a8c5253d798e896d79e875a2b6f030843ba2aacd6c772b62a19a213

Observation d58f2223-4fbb-4e7b-acb1-1aaae450cb00 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.760883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.760883Z digest=sha256:e903e101e663a7f3136232b217e01f601661e4b50db4f8f820e227aff4716ee7

Observation 261e5e7f-f591-42d2-83f9-619e9b9e5222 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Semantic understanding of scenes through the ade20k dataset

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:59.614370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:57.908944Z digest=sha256:0191af94162b0b2428c49cf7433ba6e9767060bb9fded966981ac04d6b8ec508

Observation 82e65a7f-dd6f-4645-a904-782ba0bbac15 · outbound

This paper cites Modeling context in referring expressions.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Modeling context in referring expressions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.044273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.044273Z digest=sha256:5b1f8fbe7b4aa481cdb54c6151951f019bdb2a0cd7c120f29e051b97b71812c0

Observation fd832eb5-db42-4116-a646-8ef24c2443b7 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Referitgame: Referring to objects in photographs of natural scenes

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.186762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.186762Z digest=sha256:e7df7014431f244a212287d59297ec28b2ddf6e7874ced0de1cf4e24fe187392

Observation 4a8d71f9-e49c-49d5-955e-67827cfabefc · outbound

This paper cites Decoupled Weight Decay Regularization.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Decoupled Weight Decay Regularization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.338436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.338436Z digest=sha256:c87cd16d7c3b3dbacc79265da677e95338c1fafc3d797125c66dd884f57b1ea3

Observation 739b10e3-3117-4911-8823-74b581b8c97c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Lora: Low-rank adaptation of large language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.515428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.515428Z digest=sha256:a8b2276e8d0c9db4f6482fd5e23dfee7786a2b9125fc0fb5133a834985494ef8

Observation f0aab330-2e8b-4741-83d6-0fadab5e9d1f · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.666978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.666978Z digest=sha256:75c591102baf1ffbb84c647e09277b3321454e4b9fcff132f1a33fdd630ce904

Observation e9dca6ff-306a-454e-a55a-ec518aafaed8 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Lavt: Language-aware vision transformer for referring image segmentation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:59.328307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:58.783069Z digest=sha256:36d6298e0683d3554dbaea3aaeae9d3ec9cda3f620a670a4d557ed4da95d5107

Pith citing papers

Observation 9a13edfd-1395-4bfc-9ef3-7f64a467d600 · inbound

Prompting Diffusion Models for Zero-Shot Instance Segmentation cites this paper.

Prompting Diffusion Models for Zero-Shot Instance Segmentation LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:18:40.812535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T10:43:15.465858Z digest=sha256:e0546a27c4ec2e7411c697c41b2df5a866bb44210a06670b945a537dd3ffaf1f

Observation acd73b69-4fdd-43b3-8771-fa81727d664e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:18:40.812535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d9a88ac15e29dc140ccda16fd5cd787c1f03d77b88a629f19c1782aa150505d1

Observation 1441080c-8ba2-44d6-b1a4-cc3473f12a76 · inbound

Towards an automated AI-based framework for floor plan compliance checks for residential buildings cites this paper.

Towards an automated AI-based framework for floor plan compliance checks for residential buildings LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:18:40.812535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T23:28:46.879238Z digest=sha256:8cca732d4876cb94131afb0b4b615046f2bede3ce8390bceaa82dfe4632a3976