Pith. sign in

Paper Citation Record · LEDGER

PixelLM: Pixel Reasoning with Large Multimodal Model

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2312.02228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02228 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:33.993033Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:03.024533Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26fbb55f-4266-4bfc-bb44-8f5d254b1555 · inbound

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level cites this paper.

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:46.653302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:46.653302Z digest=sha256:4993acd6c3cf75e61883c34ebdb318de9eb4edb1eae89f4d4babe479d1160c5d

Observation 8f174c74-546b-4ec3-9044-5f2bc7b5e25e · inbound

HyperSeg: Towards Universal Visual Segmentation with Large Language Model cites this paper.

HyperSeg: Towards Universal Visual Segmentation with Large Language Model PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:03:54.184388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:03:54.184388Z digest=sha256:7cb275c54de84fec0bf942864810d297b1b6a03612867e3211dc83beb09798bf

Observation 078c33a5-c77f-4ce9-a57b-a11819d7d899 · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.077636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.077636Z digest=sha256:7c3c58e8210c140a1e6d3df548339318164c0eb1331f3700d3f14f816b822496

Observation d6c41679-8549-44bd-86de-4ff915752f29 · inbound

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models cites this paper.

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:59.602850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:37:59.602850Z digest=sha256:9e5fde9a3d95a88860e4bdb1b52849ccf2add24a5c3fe3363a65dd274c14429a

Observation 86681f54-61b8-4a88-a6f5-273f4ef7fc9f · inbound

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing cites this paper.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.653877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.653877Z digest=sha256:cf173d629cb046df7e2842724244cdd349f816de9205cd91bbdd2ffc8e06bc6f

Observation 5cf9db23-a1b3-4486-9bd3-6d6313aaeca9 · inbound

Advancing Visual Large Language Model for Multi-granular Versatile Perception cites this paper.

Advancing Visual Large Language Model for Multi-granular Versatile Perception PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.107930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.107930Z digest=sha256:1ede6050d5c5e756465c10451fedae3ad6d1521e095f1454e501e0bd6372ff16

Observation c34e7079-30e6-4004-92c1-01f0c58eccf7 · inbound

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System cites this paper.

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:33.993033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:29:33.993033Z digest=sha256:38118933c98e1dd0e095dc869fedf2a0c9630ad6d13442c05d4860c152bbfc00

Observation 1a35217e-96be-4a4a-bc27-61c00f6786e8 · inbound

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes cites this paper.

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:36:34.027052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T15:35:30.549656Z digest=sha256:c49afaadb11b544fff8c4addf99790cab47823fce72a3ea0513dc75622123f80

Observation cd626809-10df-47f8-9bf5-15e7b2ba0485 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.475649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:f16f8a11f69e087b6f825440e6c1fe2b88857b7f10c393a1312da1344f32463c

Observation 592de972-91a7-4d8e-b634-46c85a1f6496 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.025941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c5896cb7585aafd5f76e627feec359bc31ce16215d9def16e7f5c25cba565213