Pith. sign in

Paper Citation Record · LEDGER

LlamaSeg: Image Segmentation via Autoregressive Mask Generation

As of 19 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 4 inbound Pith citation observations for arXiv:2505.19422.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19422 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:58.783069Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:44:05.499954Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 623c40c9-214e-4fa2-a520-3b27fb5e0b64 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.765134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.765134Z digest=sha256:709b11d82696d57ec4db9022035621b9f499556728e66329eb72c83927da9917

Observation 9dcf6b2f-59a7-42ed-a7cb-9a1c7ec7e0f2 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:05.575503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:48.914096Z digest=sha256:5c45d536f93eb3129554aa38c97d08f303ad60b7c0fe43ebbc84b59620d6f786

Observation ad4dcc7b-8d3c-4721-868a-1ac35484542d · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Emu: Generative Pretraining in Multimodality

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.060461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.060461Z digest=sha256:ddff6393f0a0a133bc1f7a693ab13787c3994584a65ff35d676b8eb60393d47e

Observation 3953b53e-9e4c-4a3f-9a3f-b5b254c81aef · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visionllm: Large language model is also an open-ended decoder for vision-centric tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.248645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.248645Z digest=sha256:3f50f855bd725c52d9f5ba3fbc61120bdac742e43fbf8fcb2a04e1dd358230c5

Observation 69a2b8a6-1ccb-4816-934a-7046e61ce2e3 · outbound

This paper cites u-llava: Unifying multi-modal tasks via large language model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation u-llava: Unifying multi-modal tasks via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:05.285437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:49.357629Z digest=sha256:a6684ea3f286c69a215ea184765e3fd0a0d71a7cafdb84975a524af0a33bdeb7

Observation f390e4d9-901a-4de4-a06d-b1b5e7ec3a95 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:05.030481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:49.508094Z digest=sha256:182906c676e340aadf47d55f935d477247020aa7bdb39aaf3c879b650c6c5c0f

Observation 8704cb23-df55-4a38-8c1d-25b18d87cb06 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.688885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.688885Z digest=sha256:7162e971dd1e0672602971faed8460054f5ec1d237c2ce0dc6b39cb16915a0df

Observation 32bea79f-30e9-455d-840f-11c34eafb280 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Lisa: Reasoning segmentation via large language model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.820249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.820249Z digest=sha256:7d97bbea158375fa790dbbf112a9f0f3212fdb8858d366f91f604801b36604c7

Observation 2c526dc8-57f1-4047-9239-89c4817eb0ef · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Glamm: Pixel grounding large multimodal model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.950170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.950170Z digest=sha256:9c6fef48e7fa4a52b05794e821ad7347972b157e22a20cdc4d17327722c217c4

Observation 5430b6d1-ba48-4a40-8fdb-096520526187 · outbound

This paper cites Text4Seg: Reimagining Image Segmentation as Text Generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Text4Seg: Reimagining Image Segmentation as Text Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:50.127833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:50.127833Z digest=sha256:3eff129f784757a5b8eb71009ee6720913a1beec88e86e34c1e07f9e32fe9c79

Observation 76f4e1fc-a9d0-4370-b774-b79e204a20a1 · outbound

This paper cites Git: Towards generalist vision transformer through universal language interface.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Git: Towards generalist vision transformer through universal language interface

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:04.723255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:50.295316Z digest=sha256:eb44111b8ff36036825dec2755d99c27222f9c9c3fdbbf649bea24a90f793163

Observation 23f8c6b0-2607-469a-abe7-e9a42461feef · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Polyformer: Referring image segmentation as sequential polygon generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:04.457501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:50.454556Z digest=sha256:f062f02cf4400556a2aa351bd8dc3212c6a0c216941e73bd974d5297ea36e8b4

Observation ce8ab02f-10ec-4b1e-97c8-8968d8748a48 · outbound

This paper cites Jack of all tasks master of many: Design- ing general-purpose coarse-to-fine vision-language model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Jack of all tasks master of many: Design- ing general-purpose coarse-to-fine vision-language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:04.159697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:50.622727Z digest=sha256:68da9f996842e90e469f2edbd3f607cdfbf91006f19b5df5fb2622bc82c45016

Observation 6f1e769f-81df-49cc-945f-c5d91ad441a1 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Gsva: Generalized segmentation via multimodal large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.877209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:50.777068Z digest=sha256:49113af159a83a657c8eda4d7171157d387f36a0ab8104ffde9f3515edaad81c

Observation a7b7b965-8d84-49f2-b43e-ed3c4c9d2b08 · outbound

This paper cites Segment anything.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Segment anything

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:50.924925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:50.924925Z digest=sha256:02aab3c62f21e7e71b9676c7e957b8c6e1a3fe82e48d25ed92e205c1fe89f2bb

Observation a694bd22-80c1-4772-98f6-cb27f8dc1ff1 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Taming transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.038586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.038586Z digest=sha256:aa92539bd76419bd2df29057aced6f3900da292c4aa5db9481fc9cc566f0fe89

Observation 38cbdd2d-192c-4889-b6a8-e42843b6f85b · outbound

This paper cites Zero-shot text-to-image generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Zero-shot text-to-image generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.218501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.218501Z digest=sha256:e505d2d590b0eeacb42a46a3cd0e6439f70c2ed9453ec142e7b6ae67f9a2b2f9

Observation 3aa89c91-a4c1-4e9d-924c-79ec84744fa5 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Generating diverse high-fidelity images with vq-vae-2

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.612443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:51.335353Z digest=sha256:cb76efc8227682378cafebbdaf5757011654a47b636d7dc724bc86750b6c00d2

Observation 2754b107-31e3-4cd9-9572-3de20abc23c6 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.457069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.457069Z digest=sha256:4741e90e0006c81ee4b7229417cdb4e066da6a9130dd6b862c354e2ddc1e5c2a

Observation df5cb931-d91c-41d7-9709-d37dccc99a3b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.591027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.591027Z digest=sha256:b7ef994ebca2d374930d81db92f7ed5fdd2ee79f16f02082afe16c12a3eab195

Observation 7bc3d74e-3129-4502-90d2-5e4a7c94b055 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.761106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.761106Z digest=sha256:ab994be7decf7af2587b3af03dd529c91785725b24fe4d123dcaa38ed1a377e9

Observation 7f91fb34-e4f6-4900-9543-f5120dcab32c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:51.877240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:51.877240Z digest=sha256:4c8ebdcbeeab7cfd4d7ccdd71e86810b62e97da034c8367e8585352ba21a9117

Observation a10450b4-5556-455a-86d8-ef18a1e8fa32 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Learning transferable visual models from natural language supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.064538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.064538Z digest=sha256:a2adf3933083c2aae6a56a3fa6bea74da339e0a48dbfd55d43aad0f010cc4f10

Observation db3d08a4-bb40-4a73-b981-1a535471611c · outbound

This paper cites Attention is all you need.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Attention is all you need

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.224782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.224782Z digest=sha256:60bd68f77c43735f9e1a5a0799d593db342af7de0d02104296ab187ea529fb8f

Observation b3c316b4-94f4-47dc-88d4-9ff1e9daa74b · outbound

This paper cites Improving language understanding by generative pre-training.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Improving language understanding by generative pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.336596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.336596Z digest=sha256:827f5a7857b58886de36fa1b9c3a325a7c9b43c91cc99bf781ee41858a9ca3ee

Observation c4218a2b-b437-4ea8-90f8-21432545ca16 · outbound

This paper cites Language models are few-shot learners.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Language models are few-shot learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.477997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.477997Z digest=sha256:729736bad0f2b1e723fa34f6d93f868676955cb8982c33f07ea658497bd29f50

Observation 28da60b8-52af-4cbb-aab5-c0c1a7597ab2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.589135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.589135Z digest=sha256:64ef6c637113e9166357e284aebcd2dacf8387c813566e03fd51f77c3ef76b73

Observation 7d3435b4-2434-4c01-b521-67d0c81911a6 · outbound

This paper cites Palm: Scaling language modeling with pathways.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Palm: Scaling language modeling with pathways

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.742201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.742201Z digest=sha256:a2b22f1c5ee16dc084df22cd42d4a1fda03271f5facb69f6f202f50f0ba4e3a4

Observation 4ffcabb1-f4f2-4be8-a4d5-82edf36b0c4f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:52.902659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:52.902659Z digest=sha256:be220c2b2ef5b94aef9843179f3883b15bbaaaf7333f7c6997579da2d9617f78

Observation 37bff697-7b0e-4354-a191-617cc4e3d316 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems , 36:34892–34916, 2023.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Visual instruction tuning.Advances in neural information processing systems , 36:34892–34916, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.079223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.079223Z digest=sha256:b44f6830676a02ee99abd3799a0c4980965e5b81aff8d6b81eec91316f5d104d

Observation f3fc05cd-94a9-4966-a4be-fef7035f58d6 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.272244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.272244Z digest=sha256:26f39d090b0f223ae11a9ab2fd87e8a2de58f64ac3ba7a454b83ccebba6affde

Observation 9c0e5efd-d244-4cbf-834a-2e22166b6fe6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.373312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.373312Z digest=sha256:cf5610d17752f09da10a022029d4a00ff4beada2e6b902ddeab06a1fd9d3340c

Observation 778a211f-b458-490b-8669-8af288f77ece · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.498373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.498373Z digest=sha256:d074f6e5e4f6be3d7186c01c590b7a475eee6b666b73025143146a33bc679a2d

Observation 67503526-5313-413c-9f3b-e9846bdf4c35 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:53.654071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:53.654071Z digest=sha256:92ea2ca58f1d55598f09544072de566eb282880d2a56db562683da4f79fa2628

Observation bc4dbb58-1eb3-4735-a636-3454e1bc0e3c · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Per-pixel classification is not all you need for semantic segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.335852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:53.841047Z digest=sha256:70aee67702271dd9907bd52d502045f54d7a9cf2744f5cdbe517305aacc9daed

Observation 19ede363-d424-4c61-add1-4794e31cb037 · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Segformer: Simple and efficient design for semantic segmentation with transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:03.067538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:53.964977Z digest=sha256:d7f2d9aa9426b98a969cb41d22dfe3f99f49fa22aa2ecff9672bceeefa22daf2

Observation 86e01068-3a53-4f58-9c6f-1d3e6d932ddd · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:02.806222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:54.105161Z digest=sha256:0a80fce4c2b361c5c83fb607495b3e2e06b7a5395626e9ca941e3d626f5377b9

Observation 7040646a-69a7-4de8-94d8-54fc8a91c106 · outbound

This paper cites Vision-language transformer and query generation for referring segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Vision-language transformer and query generation for referring segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:02.504780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:54.320451Z digest=sha256:68115f82dc0c2b1a1b16690b05c74b43a2209873f45d87c309599dce54034e99

Observation 336a06b2-7b56-4ea6-844f-9e747b18ab12 · outbound

This paper cites Cris: Clip-driven referring image segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Cris: Clip-driven referring image segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:02.165503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:54.477855Z digest=sha256:7ba60b52edca5121d54c21cae2a31cdf045f45f445055dfdf056fd44064b8012

Observation e8044630-6258-44fc-94f5-c111eaddc3cc · outbound

This paper cites Restr: Convolution- free referring image segmentation using transformers.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Restr: Convolution- free referring image segmentation using transformers

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:01.859025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:54.628990Z digest=sha256:68384aa34485423c513ef3dde70a4b4cc9105620471ee1aba563d13aeefe720b

Observation a41cf8c1-127a-4aa8-bca1-6ada93eab595 · outbound

This paper cites Gres: Generalized referring expression segmen- tation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Gres: Generalized referring expression segmen- tation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:01.579004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:54.809115Z digest=sha256:6f76cb09d9f8fc1e20ee8f2e03f0383572136f525d7ac8fde76ff2303f339e7e

Observation abd4dfea-31d3-4e2e-8979-37aecf7ebbfc · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Pixellm: Pixel reasoning with large multimodal model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:54.960384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:54.960384Z digest=sha256:f6d1deefe4b2915d86528eb774d8ec31911c41b45711ce8b003761bb82b41dc2

Observation f127ebe1-442f-4a9c-8b4f-6c0ce9a41818 · outbound

This paper cites Generative semantic segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Generative semantic segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:01.224061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:55.096922Z digest=sha256:dc01bc48889d24c8b20afd8965ab96dfd1fff30b6a74414539ea87150fe0bdc9

Observation c632f009-f28c-4ec8-a450-770c44e795a9 · outbound

This paper cites All in tokens: Unifying output space of visual tasks via soft token.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation All in tokens: Unifying output space of visual tasks via soft token

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:00.848461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:55.239331Z digest=sha256:ba7af426ff098574a92ab79a6aaba55d901c0d94c80ec50ccc2f7afe6a929ac9

Observation e685d759-ba8f-47d8-b3a2-b88bacd1eddc · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.397888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.397888Z digest=sha256:4aac0bd34cddbdd6c85e246d48ca6c6d1e77231593043fadfae79ac9e76000f0

Observation 576d1f19-6221-437a-8a5b-eeba8c509471 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:00.528506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:55.554486Z digest=sha256:9992023a9100bb874986c9fbe21ae6f16cf4eda5e8cf2356c9f015343210ecb4

Observation a7d720ca-b63c-439a-8d53-88cc3e268380 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.721329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.721329Z digest=sha256:23c1980f01bbe1599193651358bf93fe3b1ce1916ece3dfaffbe483f0aca6d98

Observation 3ce10ca2-5e5e-45ba-a977-153754d25f46 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.907967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.907967Z digest=sha256:bf7eb5b1642ba560ce186ab6dd9a510a07105770de4c5274d477577f225fb8ce

Observation c7b92e2a-86c0-4a27-aff9-bb5b3b3c847b · outbound

This paper cites Neural discrete representation learning.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Neural discrete representation learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.052079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.052079Z digest=sha256:d031082a9d306165b06352a6b1ece85cc63b4a96ca73f836f321a85be6d294b2

Observation 4d710822-ad3f-410b-b1ee-8813fcc589ec · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Vector-quantized Image Modeling with Improved VQGAN

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.238340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.238340Z digest=sha256:b85dcfac7883ce3976f6a8f013871a25a638061728f73ceb3d7c9f77605d66e4

Observation 2d1b0731-6331-444c-bf65-888d1f32a6d2 · outbound

This paper cites Movq: Modulating quantized vectors for high-fidelity image generation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Movq: Modulating quantized vectors for high-fidelity image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:00.276993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:56.342837Z digest=sha256:8579b6e115890da029d7fa9505c103fff6b50e2a537c3421a6221ab908258cf1

Observation 1ad97a92-c8cf-40fb-97b2-e858d76d3741 · outbound

This paper cites Auto-encoding variational bayes, 2013.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Auto-encoding variational bayes, 2013

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.481060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.481060Z digest=sha256:2a69760f3bd7570d700b830aa65b8b85ba0fe10c52bf52e6b918529d417b1f3d

Observation 2fbe24c8-87c0-4efc-84f2-4d69eb3d4cce · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.599570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.599570Z digest=sha256:9b765477593b4f659a740aa108917f9fd3e8f7d480b328f17c9da09e93a5d31d

Observation 7f7fb99c-e21f-425a-8010-925854e1df01 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Gaussian Error Linear Units (GELUs)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.761381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.761381Z digest=sha256:de160fd970207a73aab642a60d4586706087cb193cc95ea79fe21beaf4d61908

Observation 42735c2f-3bc0-46f5-a769-d5050f6abdcc · outbound

This paper cites Root mean square layer normalization.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Root mean square layer normalization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:56.940918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:56.940918Z digest=sha256:987905fe4e08f5395afe5fb170b286b3f11a95213ab887978b0d3c37f067ea57

Observation e693b3d5-37b3-47b8-8e2c-1867145ed260 · outbound

This paper cites GLU Variants Improve Transformer.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation GLU Variants Improve Transformer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.057555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.057555Z digest=sha256:3b04e6c793dff80cd11bff2908e7ab333ba0d2395c58be3c8b5ad53e034cb72f

Observation 406fd3ad-69bb-4710-a8f4-8fdf63e9dfd7 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Roformer: Enhanced transformer with rotary position embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.179181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.179181Z digest=sha256:a0786592891f3789a386960c60c9ab18ca27958ccf41c7063349b9cb20ef9184

Observation 3e7691ad-1c85-4d6a-8646-b65250830ad9 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.313128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.313128Z digest=sha256:c342ab51368f264ce6bafed68b06567934bc596a2183726150b8dc5a88d091ec

Observation cef5d445-2c91-408e-ae4f-e378343b352d · outbound

This paper cites Sigmoid loss for language image pre-training.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Sigmoid loss for language image pre-training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.445520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.445520Z digest=sha256:bff04b130103956730661868e7efd6abf5d78159b6ddfa1c92890259c88fbf42

Observation 1a297325-00bb-422e-8fae-146270f85ee3 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Coco-stuff: Thing and stuff classes in context

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:59.900587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:57.626054Z digest=sha256:0e55a883c34b04b89adc2de81557588465605aa3bc0935d44c2a527e7dba7c00

Observation d58f2223-4fbb-4e7b-acb1-1aaae450cb00 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:57.760883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:57.760883Z digest=sha256:b45a35785cfa114ed5b43870e8dd7130ffaa1bcedd158ebc9175daad39cee961

Observation 261e5e7f-f591-42d2-83f9-619e9b9e5222 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Semantic understanding of scenes through the ade20k dataset

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:59.614370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:57.908944Z digest=sha256:ad93c129a099772fd33dc2774ebf41d7a089084cfaba0811717c363c6d11dd5e

Observation 82e65a7f-dd6f-4645-a904-782ba0bbac15 · outbound

This paper cites Modeling context in referring expressions.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Modeling context in referring expressions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.044273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.044273Z digest=sha256:6f54451462e3f97cef08fc8fdb1ccb144decf49ff1a65323c757df061b0b49ac

Observation fd832eb5-db42-4116-a646-8ef24c2443b7 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Referitgame: Referring to objects in photographs of natural scenes

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.186762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.186762Z digest=sha256:7f5c0b9ad05829881730871e8be04fcba93a92d0d299685899fc0d2a4a03fcf4

Observation 4a8d71f9-e49c-49d5-955e-67827cfabefc · outbound

This paper cites Decoupled Weight Decay Regularization.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Decoupled Weight Decay Regularization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.338436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.338436Z digest=sha256:8f411d6a0badc636bd0a4432243b6d9bc3664b383d6039df8cfede16a40945ee

Observation 739b10e3-3117-4911-8823-74b581b8c97c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Lora: Low-rank adaptation of large language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.515428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.515428Z digest=sha256:3327a33e7f8c799b6665e6e7e06c9eaf017b4fdd2f80da86625c12859de6125c

Observation f0aab330-2e8b-4741-83d6-0fadab5e9d1f · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:58.666978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:58.666978Z digest=sha256:2305bde7d5913e8c227a0311f839455ebcc1acee46262824981b061b82af5ded

Observation e9dca6ff-306a-454e-a55a-ec518aafaed8 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Lavt: Language-aware vision transformer for referring image segmentation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:59.328307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:58.783069Z digest=sha256:e2c325a2a6224629d90e29ef69546240e08d448875643f5cd2ff5a282f2bb36d

Pith citing papers

Observation 9a13edfd-1395-4bfc-9ef3-7f64a467d600 · inbound

Prompting Diffusion Models for Zero-Shot Instance Segmentation cites this paper.

Prompting Diffusion Models for Zero-Shot Instance Segmentation LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:18:40.812535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T10:43:15.465858Z digest=sha256:d0abea05ab590f4dbadf6ce52949d2e19593545836a26b096c0cca3ce9b3a1e2

Observation acd73b69-4fdd-43b3-8771-fa81727d664e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:18:40.812535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:e15fdc97a87f1684e3725ddbdb6f3785901d99c4ab9f8f8928926ef4e5b4f48e

Observation 1441080c-8ba2-44d6-b1a4-cc3473f12a76 · inbound

Towards an automated AI-based framework for floor plan compliance checks for residential buildings cites this paper.

Towards an automated AI-based framework for floor plan compliance checks for residential buildings LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:18:40.812535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-02T23:28:46.879238Z digest=sha256:c5821e278ffde58b3c63231d14261055667592e13e81262acec9c01a32004761

Observation 67c919f1-6040-4f6a-9812-99003843668f · inbound

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment cites this paper.

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:44:05.499954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:44:05.499954Z digest=sha256:40d53e624164d0106aff4ee622d0b4fdc0fb0f6e564643d30bf5051f06271f70