Pith. sign in

Paper Citation Record · LEDGER

PixelThink: Towards Efficient Chain-of-Pixel Reasoning

As of 7 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 5 inbound Pith citation observations for arXiv:2505.23727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23727 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:49.325182Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T13:38:03.546834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.964053Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9c530e1-301e-452a-aa53-0f4f72eded95 · outbound

This paper cites Modeling context in referring expressions.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Modeling context in referring expressions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:40.528887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:40.528887Z digest=sha256:ae1223d32baeb83052fd2af2b5c0048bfd72c5e9304db84b2824f030bfd80ea5

Observation c85d2a63-db7d-4920-a030-357df615d65e · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Lisa: Reasoning segmentation via large language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:56.662142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:40.599512Z digest=sha256:6fda7065078b51ecc8940d67430d95f39b61f63e51581ab70ed889e43ed16ae7

Observation 3b7ac6dd-0d72-41c7-ad02-498abbe7684f · outbound

This paper cites POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:51.290601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:40.721525Z digest=sha256:fd25e83f11d276c46e9527b9389ac6bcbe0bbfbbc25a48878eadeea9b237c5b7

Observation abded3a6-ab1e-4a87-badf-df1d9f0098f5 · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:56.424010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:40.821394Z digest=sha256:1d047004fc7a9fa98e72d31d83f1882c3f9ed144c2f5eba3bb01cbeb3d41fe82

Observation aadcc3f0-481f-4d1a-b49f-36a5aec754c1 · outbound

This paper cites Masked- attention mask transformer for universal image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Masked- attention mask transformer for universal image segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:56.179089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:40.960493Z digest=sha256:484491e3ebf4bd835bd6d3fd5a11b6eb5073a7fdb3995287c97d2409e102bfaa

Observation e5c4088e-6bba-49d7-86a4-ff1ea31018a1 · outbound

This paper cites Mask r-cnn.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Mask r-cnn

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.940441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:41.066913Z digest=sha256:03d19f766ac37c8266ddd6a8c49ed3472fe00d665ae4e67b8ab305dd1f0d8e79

Observation f1fd865f-308e-44e0-8842-1b526f9c4c76 · outbound

This paper cites Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.Advances in Neural Information Processing Systems, 36:26650–26685, 2023.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.Advances in Neural Information Processing Systems, 36:26650–26685, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.682226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:41.202004Z digest=sha256:04fa4c1ffcf364b4ec2862bdf54a70c4a28cd748b5b50adc40437ad2dde8fb41

Observation 4f7a7ac9-3408-41b2-a512-b91ee6d9304d · outbound

This paper cites Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.330241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.330241Z digest=sha256:e97ff903ff4161e972ae60e62ae6230110fb025da1893b656e11406dd401028d

Observation 5bcc5003-2b08-4172-b713-edda78ccc6d8 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.425108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.425108Z digest=sha256:2dc50c5e318c929e1847db09ab40e13ea07537b56bd9d6d77334102b564d5b68

Observation 39bd3ae6-0ea5-48c6-a8d7-f6e4be4c5309 · outbound

This paper cites Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.564697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.564697Z digest=sha256:0b855d52cded4909dfb5cd81df21fa4073db8b82215620acd1daa3c23d54bbf8

Observation 29c39849-422d-4183-ab12-bc75d4212f85 · outbound

This paper cites Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.678292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.678292Z digest=sha256:7eecc5b416bcfd8ef412f145d7b44307c10771d3a5a25de6a5f92fef0f9311c2

Observation c14fc619-572c-43c3-9df3-d068ec168ecb · outbound

This paper cites Improved baselines with visual instruction tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.433582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:41.807075Z digest=sha256:04040bbfefcc681e1089707d1b41c0a4ec89a6259e0991ec8704a800c3860fe6

Observation 25340ec9-6150-4d27-bc9a-75ff2dc3bd1a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.894174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.894174Z digest=sha256:9e0158e77fe070307b897dc127fe3fed4e460493563c71de41499188c10c2a15

Observation 172337ab-ef93-4e18-aa86-a59beb93d021 · outbound

This paper cites Qwen2.5-VL Technical Report.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.033959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.033959Z digest=sha256:7c4b28eadf3a84631935f1fb52d417e5ef0bb3752e293529724f3e0f32aad35e

Observation 9e789f4d-afc6-42ea-8178-4bad8686c8a5 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Pixellm: Pixel reasoning with large multimodal model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.209420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:42.147050Z digest=sha256:9856bbf695920296e4377369ce53382ab00f3fbef4faa28a21f77c717b04f67e

Observation 5809c1e1-7142-4f80-bbad-4ed6e7bb8d8a · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.893435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:42.202679Z digest=sha256:ec67596b5adf5582904883919cf3bc14d3b7e15750be67d3276f142c67ff7ca7

Observation 93fe76aa-8ec5-43d7-bc3c-054e0241a974 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning One token to seg them all: Language instructed reasoning segmentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.328505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.328505Z digest=sha256:e845d09b89fe2e4b602ad979808e318f917fd64eb6633671a3ec407d893ef7b5

Observation a7f9f543-9be1-45d9-9a85-494486069e7f · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.478046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.478046Z digest=sha256:951113bc878b1b9ac7677f189a4c0bb6cb6e1d7ef07e912992dd1affd73398c2

Observation ad735f36-99db-4e88-9149-6f6265565e40 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.624891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.624891Z digest=sha256:3a48ce84147e66deaa25d4b831a1eb15c15e0f1570197ff567c4acfcdff383cb

Observation 9babf3c8-dab2-4c2a-8a6f-d0b944961067 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.775228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.775228Z digest=sha256:1ae78a10fefe6f47b5a31230974cdb11d0b210f9d7cbe21e55af9d80cc7e58a1

Observation 758ab3e0-ee7c-4a81-a153-d223a9a14567 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.916533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.916533Z digest=sha256:3813d5dcf2ad167811985e080cf92097418bf2622b65fe946b099c0380c0fc14

Observation 509240d8-4f57-4570-83e5-c62b8bc25153 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.033198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.033198Z digest=sha256:8ec3caa8451dbfbfaad0dfc7c2f330382308bad4367e8fd9038a0b248b655b88

Observation 5b48a37c-a04e-4115-b913-a8236ed149d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.144525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.144525Z digest=sha256:33714c528c66be47e83b659318a8aed5b0b42237b200e9e16e2ec43200306dc6

Observation ea4145ac-7bb6-4db7-a498-3cd8d0a6ce32 · outbound

This paper cites R1-V.https://github.com/Deep-Agent/R1-V?tab=readme-ov-file, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning R1-V.https://github.com/Deep-Agent/R1-V?tab=readme-ov-file, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.647463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:43.256104Z digest=sha256:b0b391cb8b9b47d75f6a40ff2f7fe0122c3d1301a2238351ecbbfc92bfa6973b

Observation 9498b81c-33b4-406e-8272-b15b92f26e6f · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.373987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.373987Z digest=sha256:9d8abf8fd90571d554950b3107b61f40b8387b1b26df56b19532efde7c51e639

Observation a51758b7-d85c-4ca0-9bb7-f7403779b27d · outbound

This paper cites Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.494739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.494739Z digest=sha256:5787773355a848e59dcb18d35c4fc3415f8689070a95c47793ce9b47ce527e9b

Observation 4b7ce38a-bfce-4751-a9ed-51581b000b01 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.636519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.636519Z digest=sha256:2c0274dd1534d77c0fd4a014d4f2105706a569005f41663202b3add2435750c7

Observation f5d7b6ea-9da5-4967-b817-2dc30cd3c00a · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring expression segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Sam4mllm: Enhance multi-modal large language model for referring expression segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.763499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.763499Z digest=sha256:bf6bf1208d84838dec102435c7f1cca9f42835b6fead3ecba1ca55b15a8835d0

Observation f10a5cc4-8fc3-457d-bfd3-bd3ff5b4bb19 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.891818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.891818Z digest=sha256:6f3bb5df9f153baaef4b78c4b3ef19a20c3a5c9f43c743d131553e371567d816

Observation aeb88d7b-eadd-4863-a39f-ccce7300b521 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Referitgame: Referring to objects in photographs of natural scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.420700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.013359Z digest=sha256:766040b1744f366a8bbc316c831fb3dec6b827b04dee44fb58c4ecd9f1da9bfe

Observation def141f6-31f6-4bd8-9d3f-90333b08f475 · outbound

This paper cites Towards robust referring image segmentation.IEEE Transactions on Image Processing, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Towards robust referring image segmentation.IEEE Transactions on Image Processing, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.170448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.109823Z digest=sha256:d8991838b04c8417a5355e1305a71f6cda632f25892d5e96e65ada2e860d61bb

Observation dcab0df8-cc3b-46ea-8408-b9dc3b5e632a · outbound

This paper cites Remamber: Referring image segmentation with mamba twister.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Remamber: Referring image segmentation with mamba twister

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.980694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.249632Z digest=sha256:983cf2e7bfbb200108cbfda40167eddb32e28dd42a2de6ccd0e249fe614f4cf3

Observation 7c9a332f-a64a-45dc-82a1-15e3cf5f4a24 · outbound

This paper cites Mask grounding for referring image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Mask grounding for referring image segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.369152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:44.369152Z digest=sha256:44fba8a94590e840d7467bd4e0a87c8e97550873105d6cab8be816b275f11ee1

Observation 120cba28-27d0-4f29-8a95-a4fba548ade5 · outbound

This paper cites A mutual supervision framework for referring expression segmentation and generation.International Journal of Computer Vision, pages 1–16, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning A mutual supervision framework for referring expression segmentation and generation.International Journal of Computer Vision, pages 1–16, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.751979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.489010Z digest=sha256:56ff51ba85f41fae6ea474a4ba49f3d944e7af6d82af82a85b1503bc2c302d77

Observation a12b4204-9fa7-4b6d-888b-6f2e11a24395 · outbound

This paper cites Pixel-sail: Single transformer for pixel-grounded understanding.arXiv, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Pixel-sail: Single transformer for pixel-grounded understanding.arXiv, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.553534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.593422Z digest=sha256:e662d5f0ecf7a030b041439209d70815441a758680a1884df4364244ef4948e8

Observation 4f3ef293-2857-499a-8ead-9cf4ea864dc7 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning U-net: Convolutional networks for biomedical image segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.343753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.733818Z digest=sha256:9a944af21f68bbe031503c239a837374e42137817d3575fc2f10eeea920ec810

Observation db27352c-e588-46ed-b42e-3b1aab24532b · outbound

This paper cites Segment anything.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Segment anything

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.055280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:44.872837Z digest=sha256:49023df8a25d2bd06234a9ed27a1f83dbca4fbcd33d9a76b4d0279cdb1e02e04

Observation 223c34e4-a933-4b81-a9e9-357fc0e6451b · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.966934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:44.966934Z digest=sha256:ae419307b4c8fc6c7e4e273b45f2ceac02efa64cc188a878c83bfc3efcf28278

Observation 0a77e824-65fd-4963-b8e0-73070732e6bf · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos.arXiv, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos.arXiv, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.800808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:45.065669Z digest=sha256:17877b86c3f5815006ea6a5fb051db8c938962095fb0654023f210966dbed796

Observation ce371911-473f-4dc6-9c50-f66fe4f1888f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 35:24824–24837, 2022.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 35:24824–24837, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.183443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.183443Z digest=sha256:0f15a4b408eb5b39cd2aa224e983f57a2c359136a863930add3780c6dcb22aeb

Observation c06fddbc-7fd3-4fb6-bc7f-1c763812e03a · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.327386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.327386Z digest=sha256:7a0fd6cd7ae38defd602b769bd55b2bc8310bbf2f964d4c8ca0bae8b21ea950c

Observation 953f057e-1778-438b-9544-48944fbe92ca · outbound

This paper cites Towards Better Chain-of-Thought Prompting Strategies: A Survey.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Towards Better Chain-of-Thought Prompting Strategies: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.422348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.422348Z digest=sha256:3de0cc2190da01149cd7a924f744e79f02eef2f6ceebac5e501c083ddfff1c90

Observation 18fc7829-96fa-47f1-9703-a59bde8fb76f · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.543162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.543162Z digest=sha256:ce4ae2a2d650e4c9c9275cfb460dfd06708042fe30c6c772c465583d74cc4d94

Observation 98bcf2ab-b1af-405d-b0dd-0ffed01820e9 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.679803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.679803Z digest=sha256:2e13f5001318eea36d6bed8c33166fc3ab7434471d65042517b6d1dc843fbf3b

Observation 495b41d2-2427-436a-9c92-bc82e99c77d2 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.805388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.805388Z digest=sha256:1537c18823e65690c17f32afac249d1328b31aaf76db5622a0a32a93dc19fbc2

Observation 07be56fd-eeef-47ac-88c1-94ce2181fd92 · outbound

This paper cites OpenAI o1.https://openai.com/o1/, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning OpenAI o1.https://openai.com/o1/, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.891852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.891852Z digest=sha256:7197e6dcdfbf15e31448d2f58ebb9e7164bb1c178e19737f91238b5a4e2d571c

Observation 27c9ec78-1897-4946-8a51-8a205da80bc8 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.977413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.977413Z digest=sha256:e8064201e6ab4145cde5e394dcc7354e63b851ca72bf0886e45259cf3addc98a

Observation 07d63e1e-308a-4252-b7d1-74c29b4ca781 · outbound

This paper cites s1: Simple test-time scaling.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning s1: Simple test-time scaling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.077612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.077612Z digest=sha256:49b210ebdc8d8060f7dd4044598d56803c4f7786a3db0756ad4c16b2ab82d02e

Observation e8b51122-ee9a-4f4c-a762-8a2f07fc51e9 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.144862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.144862Z digest=sha256:0b06bd990016e7bc4880dd39e4c4c4a0a86bf96709cb496d2f4722b841b0d719

Observation 28802efd-f994-49e1-ae31-14b1dbdd4de2 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TTRL: Test-Time Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.198402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.198402Z digest=sha256:350cb96878f0ec4a9f884149e9e8bad0b421d79e090c164c24628e02c5319c8c

Observation 504f63f2-81b5-433b-b5cc-aebbe956ac3b · outbound

This paper cites Open R1 Multimodal.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Open R1 Multimodal

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.576238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:46.260444Z digest=sha256:0ea3080cd119160f9dcddb3d63f4aa9cc5134c0b1af77f376c423f75cee705f4

Observation ba38a7ee-5ee6-472f-b365-1379f93f506a · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.355639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.355639Z digest=sha256:db4b152a00411e6166e91bd8b4fbeabe8d38551805c61f06ddc29e0d0733edb1

Observation c03c36a2-7818-47bc-9124-33934af2f74a · outbound

This paper cites Efficient Inference for Large Reasoning Models: A Survey.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Efficient Inference for Large Reasoning Models: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.479922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.479922Z digest=sha256:2b9a2e1074c3191af4c327b1355a5bbf222496dc49b80af6ee2b1508d1f6ff07

Observation 1d42fa26-364f-44fe-b663-9eff29dccd5e · outbound

This paper cites A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.575425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.575425Z digest=sha256:07c574de7b15d925416e7aef1b938fb36ee62941fb8394058fcbde850ab4cba3

Observation e7fa8cc1-75ad-43ec-82b9-63bea8c08157 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Token-Budget-Aware LLM Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.718839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.718839Z digest=sha256:01cb538e5e1c87e91aa589871b7d0edaec7cf6a1cd92616dd10c57f5beeb5ef5

Observation 14b0b5c8-56d1-4a7d-9f8f-3a8a885a9fb9 · outbound

This paper cites CoT-Valve: Length-Compressible Chain-of-Thought Tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.813417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.813417Z digest=sha256:744de1661d2ce602f2e56a9216cab138316c0469cf843f83971a6296641310b3

Observation b9e5165b-591a-431a-a5ed-f3d76ab2ad52 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.890731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.890731Z digest=sha256:725878ec5a7ee7dfabbddb99a2c43f63084151221d051ccfa3ea061f554f0991

Observation 2568004f-7a97-4103-a80f-b990e3e95664 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Chain of Draft: Thinking Faster by Writing Less

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.989840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.989840Z digest=sha256:1127b8e630159636f6b7a8395e87278198729b84258ab6cfb9a841ef0d25065f

Observation 42aededb-4303-48bd-8e9c-30030cd75dd7 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.103317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.103317Z digest=sha256:f508d55b99ff9dde6ccbf73f9adf7dc0e5674e3342e57670cb8c8affe426accc

Observation bc08dba0-d75b-4fc8-90db-9494c4c596c0 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.190857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.190857Z digest=sha256:a27e8927508066701ab17800229ec6ebb111929b3368aeb9fb96b6006208bf16

Observation d88c928a-1794-4b88-86e6-323e3c1d21b5 · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Self-Training Elicits Concise Reasoning in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.270161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.270161Z digest=sha256:baad1217a7f54daa06f098dfb080d6f8be35eed089c334d2c67ba614801071a5

Observation d5f147fd-2b85-48dd-9d0f-711bc0cbc900 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.369962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.369962Z digest=sha256:2429f2c0b55ddf25cee726e9ff268d7feb7aef62dd2e5e78c65a277df975726f

Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.521580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.521580Z digest=sha256:bc4bdaa06f9b04a79b01144a6220b438db33bc13e164af3cd6a09eef0a6b362d

Observation ee0d314a-5c6d-4fb8-8488-8e40fee127b2 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.643037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.643037Z digest=sha256:3f57d571559cad65b08a1ebf6307060bea108b9508409d20c8960217c83c7afd

Observation a8b5ad51-a5eb-4969-9f15-cc7d751afe62 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.arXiv preprint arXiv:2412.04467, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Visionzip: Longer is better but not necessary in vision language models.arXiv preprint arXiv:2412.04467, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.770684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.770684Z digest=sha256:b6612ccbf292231c15be2d3eef4ca3b5cd570b1b5d7949a28eb7b0f7a3743ea0

Observation fc00c54a-baa8-4102-a00c-f34c8b0ac5ab · outbound

This paper cites Learning to Inference Adaptively for Multimodal Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Learning to Inference Adaptively for Multimodal Large Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:50.149770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:47.885550Z digest=sha256:f887c968012ff11669125b3063e07cbe73ebed1f85d82b83cf63b6ee6a5ee8cb

Observation 05fd2193-04e6-4edf-87ed-77bfb4816150 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning SAM 2: Segment Anything in Images and Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.987803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.987803Z digest=sha256:f773289f4566268d8e6b4d8e5ca65e371d8ecb70e9507f729c0c51a86c4a3f6a

Observation e9080b37-d52e-497f-a384-d23c3b099d9b · outbound

This paper cites Minimum-Margin Active Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Minimum-Margin Active Learning

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:49.840522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:48.090813Z digest=sha256:2cc046f2d72eaaf25ab76bc55d1694e16b3515d22e8fe7a6afec4dd43df18f66

Observation 55a21d8a-69df-46c6-849d-d76362d7ac9e · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Chain-of-Thought Reasoning Without Prompting

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.207589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:48.207589Z digest=sha256:07a9ea9a27ec16ba1a4020aed79deaf7338a76708edb2118e41e2191d8c474d2

Observation f1e504b5-8e26-417b-986c-6ce42b072558 · outbound

This paper cites Qwen2.5 Technical Report.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Qwen2.5 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.336602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:48.336602Z digest=sha256:fd6a105baf1fce18a97be716313490b48d49f41ed38c1e51fc1010c1dfa27ebb

Observation 0bd56ddb-6426-4bf2-a465-792667874fef · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.361036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:48.453340Z digest=sha256:dac973d7cbfead12650728fd7902dc05b85636299b31454ae9f40b85456884c3

Observation 3e48d4c8-304e-443f-a12c-9202b5c50058 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Open-vocabulary semantic segmentation with mask-adapted clip

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.155138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:48.580296Z digest=sha256:319b796e219bb4f45cb3c09bac4e0c6d169103fc7f5adf9773da5d71478e228b

Observation cd92b246-cc7e-46fa-96a7-300299d3d122 · outbound

This paper cites Gres: Generalized referring expression segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Gres: Generalized referring expression segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:51.959281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:48.721713Z digest=sha256:c46540dae814d3db889943a4a353593a6f2f6e677108c5cc0b1a7ba09b1d0d4d

Observation c89fc2ea-7041-4cd5-b4f4-0305f4e92a61 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.814103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:48.814103Z digest=sha256:795a03d64414196b18f05d3f3c1f3a075eb074c35d202fde1c7123a1d6602e54

Observation 92fdd583-08e4-482b-ba24-7947dd364e63 · outbound

This paper cites Lavt: Language- aware vision transformer for referring image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Lavt: Language- aware vision transformer for referring image segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:51.710334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:48.921057Z digest=sha256:b18ef28c6a78982e3c9ab4d370dcdb5639cc8c2d0bfdb4d22d1a0d39c25feb08

Observation ce38aa49-aec0-4121-b092-3fff4802a0c3 · outbound

This paper cites Perceptiongpt: Effectively fusing visual perception into llm.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Perceptiongpt: Effectively fusing visual perception into llm

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:51.493970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:48.978262Z digest=sha256:d0a03e2c0782b71a827b1420da4d38b67149485c0040423723b455c5901a5a6d

Observation 15688ab7-2a7c-4dd4-b0ac-c2039a0c824a · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.arXiv preprint arXiv:2503.16188, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.arXiv preprint arXiv:2503.16188, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.027105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.027105Z digest=sha256:530fd5702b2e14c9eca975cfa13fa08250e51650383dc60749ed94e487a4cbf5

Observation d9f8daec-e419-4683-a770-7aee54f155c4 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.097448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.097448Z digest=sha256:6b83fd56fac02b8ecf8fcf29962f7fda642e9c519b8b17f52c6962386bd2b3d1

Observation 3146d7b9-8f90-4442-8c82-8e85a243b33e · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.152191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.152191Z digest=sha256:0a83905800c5f4b670be6f52c1cc31b9c3836044017fc305f96885a8c6d6820c

Observation 8d47115b-baf7-4cf6-a4dc-632836c4a70c · outbound

This paper cites Microsoft coco: Common objects in context.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Microsoft coco: Common objects in context

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.226688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.226688Z digest=sha256:a2b66aefcfda299167e07864d3d1ef7af654dc8b4c28f32d46cdea6eed27b411

Observation a459d7db-a80b-4667-86b6-87f82ee0a03a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.325182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.325182Z digest=sha256:82bda583c17de2d74833188d40b9def6baf8ab12c20be88704b146d3ebc522cf

Pith citing papers

Observation 65e3b23c-75ec-4b0c-b5cb-825c7e0ef0d9 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:06.680384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:e4f65f7e8dff355cbd40b5ca4bc97508467fb4e7e98e336e352c8ff0d48b8ee6

Observation 2aff3429-372d-4715-8e88-22264670b1a0 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.842628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:3ddd0943b2695e29420120a94b89097b4067b06b5bfbc2716a655142b10dcc79

Observation eef658b0-ee47-4ffd-a49d-4c206216bc64 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.966001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d303562b1900b6d11a13a8dae1fe3e3d5ae0e3ae703f5365cedb999be247c63e

Observation a391c2ba-5640-49db-8efe-ca035b958005 · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.636192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:e7e99687b35f26aee52307f71ad17ff37c51a3e3acb11c8c7201f881967d9a6f

Observation 93f3c467-b83d-4ba6-b091-288392f93fe6 · inbound

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation cites this paper.

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T13:38:03.546834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T13:38:03.546834Z digest=sha256:a46ee1e6df78124182553756d0165b828c3b5b5da0c521433068f792416656fc