Pith. sign in

Paper Citation Record · LEDGER

PerceptionGPT: Effectively Fusing Visual Perception into LLM

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2311.06612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.06612 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:18:52.509837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:36:34.000271Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1aed31e-0ae5-4527-b2e6-cb85bd7ba953 · inbound

HyperSeg: Towards Universal Visual Segmentation with Large Language Model cites this paper.

HyperSeg: Towards Universal Visual Segmentation with Large Language Model PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:03:54.171876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:03:54.171876Z digest=sha256:dbe2c0011b3619b1285ebf18ef3c1cf23f959c5eec5f1b9d49181dea28dd1f2c

Observation 2d413e29-146c-49da-a752-f51e5352264d · inbound

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models cites this paper.

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:59.594207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:37:59.594207Z digest=sha256:bbc07e56ff84d15bbb310123ed516d9a322609042ce60f2422bf9a7955a9f71b

Observation 554d3c14-77df-4a4f-b192-592bbfbc9244 · inbound

MR. Judge: Multimodal Reasoner as a Judge cites this paper.

MR. Judge: Multimodal Reasoner as a Judge PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.509837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.509837Z digest=sha256:289e99587c1584d881df8a1baac60006779b88dde58f3bd048f4b05540975189

Observation ac7378b4-d987-40b7-aad6-4158d99f5e3e · inbound

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model cites this paper.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.222475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.222475Z digest=sha256:9698725dab8015c650408eae45447897619d0cf663a0fbe0f5429e6424ea9ce4

Observation 18f123f2-0628-49c6-89d8-8e1ca92ee7d6 · inbound

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System cites this paper.

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:33.989125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:29:33.989125Z digest=sha256:3962ee70c59c5f62e32db791d98ba7347940c8ea0c0842efc450845ce0a86c2f

Observation 3b47eeb7-b1f4-4a4e-9041-6258a3ec5da0 · inbound

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes cites this paper.

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:36:34.003208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T15:35:30.549656Z digest=sha256:a9668250a2df38d1865f764b3f6a93ceaadeb663fd08a162da223b259030e256