Pith. sign in

Paper Citation Record · LEDGER

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2504.10465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10465 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.273762Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:18.024069Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d36d20c-b0ee-41cd-bbc2-03c20d0ca365 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:39:22.544748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:662d9f12fa84626bcf94fd0a9a1837dd101182934113c6199dbe1851194cbde3

Observation d43770c4-9261-42b9-901c-4092cd57a8e4 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.749504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:8f621b72413d3dbb8e5d80402ef88b72ea92ad23192c1302ed97563272860341

Observation 5984fb9c-150c-4bbd-9f4c-cb0289374dcf · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.273762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.273762Z digest=sha256:4aaf242dbbeafc7d6317bca6aebb4f4b8605944245df8c5af637a3fabbc750dd

Observation b399ce0c-bec0-4a01-8791-97bb3e3cc2dc · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:55.688403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:55.688403Z digest=sha256:190b7925cb62ff77b806e0edccc4b9488cc9b363fe289551e36868ccb3c31c06

Observation 773c2860-b542-4f6e-a3d2-2c71c3acce28 · inbound

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning cites this paper.

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:05:07.431761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:05:07.431761Z digest=sha256:2e57d562908798678dced521bcdcd3ac74a212e9c31352ee4084195a687e4ea2

Observation 925c7ed3-2ea6-488a-9936-4a04bdc85585 · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.751792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.751792Z digest=sha256:dcbca29db646fe12fc658b379bb31b0195898ea7d19a653f4bdc208a892e689a

Observation 5d347ad1-9a51-40ec-851f-1a87105d9ce2 · inbound

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs cites this paper.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.464562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.464562Z digest=sha256:322d5d2cf258bd6d60b39171a5996a384d4ab9f79967b39925c4cea49fff731e

Observation 926876d0-961d-4cb2-a56b-95a4b776275c · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:13:16.235635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:3b356e3dfb25a7ed992cae269b55c5b64931316b8f104d9908371fe2d8820fae

Observation cf4aeb30-1799-4ff3-a2a5-7e0bbab87631 · inbound

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models cites this paper.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.026551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:e70f0227e9fc91e7d0bd69703228ca55ffe94e5f623056c5d035a134699f5022

Observation d5f801c7-3793-4c59-998c-4ebfce3039e8 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:1aa397e21af295a408b49a203dfe2803a55b8b1b1cb04150b4abeead7ad246fd