Pith. sign in

Paper Citation Record · LEDGER

Object-Centric Vision Token Pruning for Vision Language Models

As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2511.20439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.20439 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:19:57.764281Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bdb75b9-0d88-489e-adf9-e36e7ed32e77 · outbound

This paper cites Qwen2.5-VL Technical Report.

Object-Centric Vision Token Pruning for Vision Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.638438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.638438Z digest=sha256:361dddc0edcc51a8d9e9bed5f05b6aa2365ee2977dd7e67982da6cf909bf665c

Observation 4f0d3d5d-0777-4283-9fe4-89b8b9592b11 · outbound

This paper cites Invariant Slot Attention: Object Discovery with Slot- Centric Reference Frames.

Object-Centric Vision Token Pruning for Vision Language Models Invariant Slot Attention: Object Discovery with Slot- Centric Reference Frames

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.643279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.643279Z digest=sha256:37388e7fc9205c3221b6efdc99a3f2900d93e00f1d61ed19ca3d5dadb7bffd27

Observation 76d49dd9-2e87-49fb-a2ea-e98103c4d516 · outbound

This paper cites Token merging: Your vit but faster.

Object-Centric Vision Token Pruning for Vision Language Models Token merging: Your vit but faster

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.647003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.647003Z digest=sha256:960308e75660bd9987d47858156233806f2989849cc312f9ca8b6e76c1b30412

Observation a3efce13-acb8-4973-aafb-ba9f97e94bbf · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Object-Centric Vision Token Pruning for Vision Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.650849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.650849Z digest=sha256:58c3001c8634e3d1b57052c6b65e311c4d6b725c80681d0b6d3900dbf313f1a9

Observation e8217cb1-0a4a-49f7-9ed5-c678a3b68b78 · outbound

This paper cites Mme: A comprehensive evaluation bench- mark for multimodal large language models.

Object-Centric Vision Token Pruning for Vision Language Models Mme: A comprehensive evaluation bench- mark for multimodal large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.654488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.654488Z digest=sha256:4b7cdcbb2d13b38217ed861d132275bbd4625bb025c758564f4bb714c1398fd0

Observation 31db8585-1472-4072-b029-37f2c5229368 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Object-Centric Vision Token Pruning for Vision Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.657702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.657702Z digest=sha256:e32a14ff7d2ea63168374478ffae99d9ed05abae96621a7c64a6ceb88bc665dc

Observation 609920c9-3a26-4a6e-a95e-ef91aa253854 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Object-Centric Vision Token Pruning for Vision Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.661068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.661068Z digest=sha256:b438630112ecc01b829be67c8f74355ae980ee5e3370d6e38b99f8f13d1c0c85

Observation 663eec7e-fe31-49d6-a943-7b3e1687816f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Object-Centric Vision Token Pruning for Vision Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.664289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.664289Z digest=sha256:05a9034a56714c52c506993cb02468a96e03811ca103ad2abfe31a3697a8fe35

Observation 54797207-daea-4a8f-ab57-2a0b46532e1f · outbound

This paper cites Improving Object-centric Learning with Query Optimization.

Object-Centric Vision Token Pruning for Vision Language Models Improving Object-centric Learning with Query Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.667395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.667395Z digest=sha256:0e28b9cdfcca2dde4c50351586fa986c79fdad5b32e6f4039f86742af6ab745b

Observation e1d3a1af-e38b-4fa0-9683-b2c3d9861204 · outbound

This paper cites Spot: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers.

Object-Centric Vision Token Pruning for Vision Language Models Spot: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.670107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.670107Z digest=sha256:d1ca71d05276c672543d5727fb3e6d0126efdc9f23bafcddd803a901b5b21365

Observation 29e27677-42f9-4bd0-a46a-886eb9446861 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Object-Centric Vision Token Pruning for Vision Language Models Seed-bench: Bench- marking multimodal large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.673071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.673071Z digest=sha256:13af0e7eeb89c7c6cef30465646720a058eba058792e7480ea8b293c0593c20a

Observation 9f61697f-ab6b-404a-8f56-d7c1bd16f01a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Object-Centric Vision Token Pruning for Vision Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.675843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.675843Z digest=sha256:97f7944a28070a81de0c967247d5afb7ef2989577ca08fc829ecb7b0c321bb4b

Observation 8648c6b5-3eec-4c1d-ba4a-0156d6f1239a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Object-Centric Vision Token Pruning for Vision Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.679008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.679008Z digest=sha256:e27689a2d501f0428015665dfc030e5e4cf27db02422bb981d171eda0845818f

Observation 781f2851-0b89-49b7-af6b-9cdee1c13a93 · outbound

This paper cites Microsoft coco: Common objects in context.

Object-Centric Vision Token Pruning for Vision Language Models Microsoft coco: Common objects in context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.682081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.682081Z digest=sha256:b34ea516a98d972308e327df83fc727e4e75c28928b8f9071300e8162f3113bd

Observation fc242561-69c5-40b5-9219-65e5fd4cfde7 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Object-Centric Vision Token Pruning for Vision Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.685000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.685000Z digest=sha256:baaa6488275c855e457e49f92a92f2ee56ce58d25f4c4c58c401094d7f921787

Observation d675fbc9-106a-4e70-b47b-0b320e632b6d · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Object-Centric Vision Token Pruning for Vision Language Models Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.687748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.687748Z digest=sha256:5774e4332c7f3cd67e88a2a69cdacacdc7ac69a040174cc2e492d93e8e5cd77b

Observation 2701cd67-e999-4c85-95a1-d10392917ddd · outbound

This paper cites HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models.

Object-Centric Vision Token Pruning for Vision Language Models HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.690392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.690392Z digest=sha256:a8230cddcd482c7be242096ae47b2e76e8e0da1841fd58d250825e0eb7078039

Observation 15fe81ce-0ad1-4070-b11a-0052199e36e1 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Object-Centric Vision Token Pruning for Vision Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.693646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.693646Z digest=sha256:9e3c83df9869194f2daa1357e9d3360bd8bf510c8a6c079edb892e40940f124e

Observation ff3236da-a8f7-46d3-88c0-312359f14136 · outbound

This paper cites Object- centric learning with slot attention.Advances in neural in- formation processing systems, 33:11525–11538, 2020.

Object-Centric Vision Token Pruning for Vision Language Models Object- centric learning with slot attention.Advances in neural in- formation processing systems, 33:11525–11538, 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.696316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.696316Z digest=sha256:b3971d527ed174312e8abd2e1e9fa8f8b7f2e7ac8c25a17530c9694330edcf4e

Observation 87a856d0-8a55-40b1-ae5f-547db0fc364b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Object-Centric Vision Token Pruning for Vision Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.699067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.699067Z digest=sha256:a35c2c870b152e66e7952b2dc737657c1902f039b27dcca3f970d9d10fb44418

Observation 44d11a48-ed1b-4212-8026-ed04f7018e8b · outbound

This paper cites Temporally consistent object-centric learning by contrasting slots.

Object-Centric Vision Token Pruning for Vision Language Models Temporally consistent object-centric learning by contrasting slots

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.702079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.702079Z digest=sha256:b046d0a6a0fe04f7c88c98948b38f4df93214be3532a696205cef16268c83eee

Observation cfb0f784-a140-4f1c-9729-15bada9374dd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Object-Centric Vision Token Pruning for Vision Language Models Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.704723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.704723Z digest=sha256:76fd5805b24d8e498fff69827f0ad961f106a7320085ed455d7358a0d3b43a2c

Observation a99ef534-51d8-492f-af74-ebadcf5362c6 · outbound

This paper cites Bridging the gap to real-world object-centric learning.

Object-Centric Vision Token Pruning for Vision Language Models Bridging the gap to real-world object-centric learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.708009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.708009Z digest=sha256:63d0e780fd635347692e7dd190e708b3411ff49b8d2979d22b6f5bab4455207b

Observation bd5e67b3-50e3-4a13-bdc5-4586ec7cc3b9 · outbound

This paper cites Towards vqa models that can read.

Object-Centric Vision Token Pruning for Vision Language Models Towards vqa models that can read

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.711377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.711377Z digest=sha256:22a5e8b1bb93458504be03ee4f138c5a68676a3791ccf2bd084e1615cbe3dac6

Observation e037fa8f-5851-4b95-ad99-51695ad72b50 · outbound

This paper cites Less is more: A sim- ple yet effective token reduction method for efficient multi- modal llms.

Object-Centric Vision Token Pruning for Vision Language Models Less is more: A sim- ple yet effective token reduction method for efficient multi- modal llms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.714392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.714392Z digest=sha256:e2125ef5d9a864bc83de5e26272de807612390bae6233a34e961c47ac84e1c01

Observation 12614eef-ba43-45c0-a693-6ef5fa328563 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,.

Object-Centric Vision Token Pruning for Vision Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.717592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.717592Z digest=sha256:fd980e7c53bf9dfbe3bcfc523816818f8766fa71b540ee3955bfd5b2acfa86b9

Observation 8744c380-d06b-41a7-81a0-c22e11551ea8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Object-Centric Vision Token Pruning for Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.721567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.721567Z digest=sha256:eb311e6f3c81c0bd00e003e7ca73cd373b9d0d715910cdcfb2fb8bcb602b3602

Observation ad154c0f-9850-42c5-bdce-09a2c25ee2a1 · outbound

This paper cites SlotDiffusion: Object-Centric Generative Mod- eling with Diffusion Models.Advances in Neural Informa- tion Processing Systems, 36:50932–50958, 2023.

Object-Centric Vision Token Pruning for Vision Language Models SlotDiffusion: Object-Centric Generative Mod- eling with Diffusion Models.Advances in Neural Informa- tion Processing Systems, 36:50932–50958, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.727020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.727020Z digest=sha256:3ae586e6a1e4cb161f511001297729ff263a74feda41734926e69cbc5afe779e

Observation 61340193-c238-4791-8e28-0eb470e7efe9 · outbound

This paper cites Conical visual concentration for efficient large vision-language models.

Object-Centric Vision Token Pruning for Vision Language Models Conical visual concentration for efficient large vision-language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.731279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.731279Z digest=sha256:7b7c3387ff70a7d7fa6dd56aa4bcf4d1c26f4d3adb3d601f9e51cfeda7eab594

Observation c0f215aa-3b7b-4206-b922-b7aaca170ec1 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Object-Centric Vision Token Pruning for Vision Language Models Visionzip: Longer is better but not necessary in vision language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.734872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.734872Z digest=sha256:6e60d6af9f87af3f5b017cdc97426d86819d2df6e601ee702ef00332fa4c8690

Observation 501c6f9c-7aad-4093-845b-4ac155d10106 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Object-Centric Vision Token Pruning for Vision Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.738313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.738313Z digest=sha256:86985cbf715de609df0fb466a6d6e849b445cc2629ea1f2d460e92e4ff0bdd46

Observation 84f066d8-5d1c-4deb-a718-97079d7bd7df · outbound

This paper cites Object-Centric Learning for Real-World Videos by Pre- dicting Temporal Feature Similarities.Advances in Neural Information Processing Systems, 36, 2024.

Object-Centric Vision Token Pruning for Vision Language Models Object-Centric Learning for Real-World Videos by Pre- dicting Temporal Feature Similarities.Advances in Neural Information Processing Systems, 36, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.741594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.741594Z digest=sha256:aad9e0279823e331335d127a00ff94aef21a7f7a144a3b70bd02058a3e02cf3f

Observation 22b0e0e0-5ed3-4dbe-863b-fbb411463446 · outbound

This paper cites Lmms-eval: Re- ality check on the evaluation of large multimodal models.

Object-Centric Vision Token Pruning for Vision Language Models Lmms-eval: Re- ality check on the evaluation of large multimodal models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.744475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.744475Z digest=sha256:5d0d3a06b8ca41abbaadcd8a10c311dd4bdc146ae09d828f2db8217b0845b1cb

Observation 41fcc725-f991-4930-89b4-fe565ca1e56d · outbound

This paper cites Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference.

Object-Centric Vision Token Pruning for Vision Language Models Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.747622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.747622Z digest=sha256:0ef68eccdf2f814151a19d36435150bb6daba48152c1af19918795e65aad0416

Observation e4d2496a-9011-448c-bea4-4daf8586d8d1 · outbound

This paper cites Predicting video slot attention queries from random slot-feature pairs.arXiv preprint arXiv:2508.22772, 2025.

Object-Centric Vision Token Pruning for Vision Language Models Predicting video slot attention queries from random slot-feature pairs.arXiv preprint arXiv:2508.22772, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.751167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.751167Z digest=sha256:56b39e1340a5c6c03e979399a41c8a762778de965491697cd520a9dcf88469b0

Observation 15f17c84-6894-4ad7-a2af-1e1a4649d094 · outbound

This paper cites Vector-Quantized Vision Foundation Model for Object-Centric Learning.

Object-Centric Vision Token Pruning for Vision Language Models Vector-Quantized Vision Foundation Model for Object-Centric Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.754330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.754330Z digest=sha256:00523f22564e9720fb3f45545e663ca0a1d128dd74fba72783a3e12397e432e2

Observation 66de32da-b792-4fa5-97a9-c7195582385b · outbound

This paper cites Smoothing Slot Attention Iterations and Recurrences.

Object-Centric Vision Token Pruning for Vision Language Models Smoothing Slot Attention Iterations and Recurrences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.757841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.757841Z digest=sha256:9e8beb64ca505ea3483f5a2b8d277030f332662c2f82668c115802b4de4b21ce

Observation cc776f21-cf19-4a1e-a9fb-260e27c90d8d · outbound

This paper cites Slot Attention with Re-Initialization and Self-Distillation.

Object-Centric Vision Token Pruning for Vision Language Models Slot Attention with Re-Initialization and Self-Distillation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.761323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.761323Z digest=sha256:1e7db10ce52ba087c717bba409c04a7d371ac843930669d1ce93ac36d7da3e19

Observation 8dea26c4-5443-416e-9183-a419fbf7c9e1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Object-Centric Vision Token Pruning for Vision Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.764281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.764281Z digest=sha256:7801dbe9e9360e55f883d8e11fe681db783fe3698777621ac836d8f8d798a6d3

Pith citing papers

No inbound Pith citation observations are available.