Pith. sign in

Paper Citation Record · LEDGER

Pix2seq: A Language Modeling Framework for Object Detection

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2109.10852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.10852 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:13:13.832416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.509967Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2bc0035-cad0-45d2-9898-aa2b889484f4 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Pix2seq: A Language Modeling Framework for Object Detection

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.060099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:975b088cd79ebaccbbb11d267899c82e3696f1458583f4fffcb3d700a2cb823d

Observation f64e0af2-a489-4db7-8691-5181f8745e60 · inbound

Kosmos-2: Grounding Multimodal Large Language Models to the World cites this paper.

Kosmos-2: Grounding Multimodal Large Language Models to the World Pix2seq: A Language Modeling Framework for Object Detection

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:19:47.994723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:19:47.907355Z digest=sha256:8d8e7c8f59f225d051302ebb4e0cd5894ed5872e78d492c6560fc796b430da9f

Observation 8983684b-913a-40e2-a290-e8fafe5acfde · inbound

GPT-Driver: Learning to Drive with GPT cites this paper.

GPT-Driver: Learning to Drive with GPT Pix2seq: A Language Modeling Framework for Object Detection

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:05:32.021160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:05:31.928650Z digest=sha256:200c0e7b682e12b96e2e6fa71d74f372732cd164967983e18d1dbb2ac46b1fe4

Observation b5e2380d-d81a-4911-813b-a6c92648deb8 · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents Pix2seq: A Language Modeling Framework for Object Detection

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.513448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:cf151d365692153f414bbfdc421ab28d8123b84bb52430704a8d995b6f549e68

Observation dd8a690e-70a4-437f-85bd-de50ae779a77 · inbound

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks cites this paper.

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks Pix2seq: A Language Modeling Framework for Object Detection

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:20:16.170336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T06:20:15.656356Z digest=sha256:30a00a909e17b0c527c8487d9c69c55def84c20b765a86a6f79836f2ab055835

Observation 7b6af6b5-3c5b-452f-bc2f-01a6769cbc49 · inbound

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation cites this paper.

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation Pix2seq: A Language Modeling Framework for Object Detection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.047201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T16:12:25.980853Z digest=sha256:81d5a298e749a837d4b00a5b034a4e677d7a27a19a706effa996328cbc78b8ad

Observation 0c20d5ca-39b0-434b-9b4d-caca9c4e1031 · inbound

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents cites this paper.

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents Pix2seq: A Language Modeling Framework for Object Detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T23:13:13.832416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:13:13.832416Z digest=sha256:945902986dd3bccce8173b1d32b83e18c9fb7192a43456375a6984b80442cd8d

Observation 5b769fe4-4d07-428c-a9d5-8968fd32bbb5 · inbound

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis cites this paper.

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis Pix2seq: A Language Modeling Framework for Object Detection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:18:18.974768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:18:18.974768Z digest=sha256:7de0a26565f43fdccc8bafe8e9b26e92f301643c85e3545b776191b96c7d95cb

Observation b3f473d8-603e-4ab5-8bae-861db00c5167 · inbound

Topo2Seq: Enhanced Topology Reasoning via Topology Sequence Learning cites this paper.

Topo2Seq: Enhanced Topology Reasoning via Topology Sequence Learning Pix2seq: A Language Modeling Framework for Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:07:35.502973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:07:35.502973Z digest=sha256:abcf9b49910a128f280f4e7fe3e5415f4486934cf3c4bcac0db841bc58d608c5

Observation 7695381b-dbd6-445b-acb7-586bd352201e · inbound

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion cites this paper.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Pix2seq: A Language Modeling Framework for Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.630304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.630304Z digest=sha256:6e95c974b5eba42066255e837fb84131432cf0bcf9de245c5c156c184ff86eac

Observation f4ebd5fc-10b3-48ac-9858-681f8a890aab · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Pix2seq: A Language Modeling Framework for Object Detection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.624177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.624177Z digest=sha256:f1f3651703f38267e1f09c0f62c3a4ab7b193276a33bc3f8146e44fc87150d34

Observation 1ed01301-9445-4456-81db-66c82418bb95 · inbound

ParkFormer: A Transformer-Based Parking Policy with Goal Embedding and Pedestrian-Aware Control cites this paper.

ParkFormer: A Transformer-Based Parking Policy with Goal Embedding and Pedestrian-Aware Control Pix2seq: A Language Modeling Framework for Object Detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:22.232547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:22.232547Z digest=sha256:ff9a565cfff6d2852c4ce27ff6afb63c21c02dff092898917f4eeb63b53f2fb5

Observation e6316a38-2a01-46b3-ba27-c207979b9568 · inbound

Style Transfer: A Decade Survey cites this paper.

Style Transfer: A Decade Survey Pix2seq: A Language Modeling Framework for Object Detection

Reference 259

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:27.661151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:27.661151Z digest=sha256:f099a501e5d44dd67b7fea9e0679a27cf46dba3d502f651a24bc6f1083e6dddd

Observation 38abb5b1-7265-4560-99b6-0bb37bca9d5c · inbound

SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions cites this paper.

SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions Pix2seq: A Language Modeling Framework for Object Detection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:44.541549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:44.541549Z digest=sha256:7d21aa8fc5fc7e67b6f94b06081b5193e52c43f9f1d9e6f7974173f757245d7c

Observation ae263cd5-40e8-4095-b55b-521668b566e7 · inbound

Parameterized Diffusion Optimization enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading cites this paper.

Parameterized Diffusion Optimization enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading Pix2seq: A Language Modeling Framework for Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:15.056216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:15.056216Z digest=sha256:2d70e21a83b605b60b41719d8013af3df81d756770b246705f13a9ca22f0638f

Observation e938ce05-d126-40cc-b325-8a904404d84e · inbound

From Provable Correctness to Probabilistic Generation: A Comparative Review of Program Synthesis Paradigms cites this paper.

From Provable Correctness to Probabilistic Generation: A Comparative Review of Program Synthesis Paradigms Pix2seq: A Language Modeling Framework for Object Detection

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:28.819122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:28.819122Z digest=sha256:5e822505a94ebfef3f9c3775b3fffb25fcf812f7990c352b1505100e74123f6c

Observation 6e42424b-8660-4faa-8e99-a0a2f5d10226 · inbound

DriveQA: Passing the Driving Knowledge Test cites this paper.

DriveQA: Passing the Driving Knowledge Test Pix2seq: A Language Modeling Framework for Object Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:58:05.084880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:58:05.084880Z digest=sha256:837b039874454d8d34b6baeff6e53cd80f4ec3ba4c76e514f72fe74bb664028d

Observation ba5068a3-5aaf-420b-90f0-38c5ea1dae20 · inbound

SAM 2++: Tracking Anything at Any Granularity cites this paper.

SAM 2++: Tracking Anything at Any Granularity Pix2seq: A Language Modeling Framework for Object Detection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:54:20.148463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T19:51:56.047515Z digest=sha256:d0e9afabf4561c4e3a5e20920374fcd423dc3d3bdc77903ef85d1901abc62d5d

Observation 15121c37-d0da-49a1-beba-ad9b4eb35cab · inbound

Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction cites this paper.

Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction Pix2seq: A Language Modeling Framework for Object Detection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:20:39.247034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:19:11.066723Z digest=sha256:191cbeff85ef244052f1753a921a14b8795aa79234f6ef9f5019667fe57c274e

Observation b24b4937-78e6-4a7f-b973-f6ea9009dc61 · inbound

Moondream Segmentation: From Words to Masks cites this paper.

Moondream Segmentation: From Words to Masks Pix2seq: A Language Modeling Framework for Object Detection

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:18:13.726512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:14:45.629804Z digest=sha256:254dadda84509f8453ae4d8e4f69857ef9382a20960efa7bfe8ebc5fe6951d76

Observation d457d7e0-47e5-4d30-917a-fbac837d9dfd · inbound

TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction cites this paper.

TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction Pix2seq: A Language Modeling Framework for Object Detection

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:30:58.362347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:37:44.436873Z digest=sha256:2e500f3327ff4ee450690eeb39a77b459a517c2622b7165a4ef981ec665fe480

Observation 8caedca2-b4bf-4222-b853-4cf29b8abe56 · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Pix2seq: A Language Modeling Framework for Object Detection

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.158931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:eac8484fce205bee7a9b82ef447836e1133d3607624364ec7afdc94127403e8f

Observation 0f3ea96d-766d-4ff2-b6fe-5f803b2ac123 · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Pix2seq: A Language Modeling Framework for Object Detection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:44:03.328300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:f67a46fff3f24ae9f060e7d131cd52bbb2c5945a8065df71d163282f69208c0e

Observation 69a5ac09-a4ce-4260-8892-aa5fe563d0e6 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding Pix2seq: A Language Modeling Framework for Object Detection

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.904809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:aa8cc820abba942af2efb47a761da3e8f17d387ff7eee374871d7882d76e8921

Observation 7acf2eb3-5cee-41dc-bae7-1311d46255e9 · inbound

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs cites this paper.

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs Pix2seq: A Language Modeling Framework for Object Detection

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.311197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:01:14.023587Z digest=sha256:f93386403c1d3f7a37d7bdf85292e20b94a5cb864012d3a68b85e4a8952d71c3

Observation 1459c1ca-0239-47d2-8ea5-58544990856c · inbound

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments cites this paper.

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments Pix2seq: A Language Modeling Framework for Object Detection

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:57.511404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:35:32.041041Z digest=sha256:fd4a8fe7c82733cf21fb1ab4a264c25764106deaea5c40f7d541f2fe2703985e

Observation ce44e7aa-d473-4b4c-950b-afbec0d58511 · inbound

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments cites this paper.

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments Pix2seq: A Language Modeling Framework for Object Detection

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:29.063968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:52:18.115419Z digest=sha256:bd01402dd4dfd8663d19604ea125e2171eae64149fd2ea270c5eb1345de427a1

Observation 679ce7b8-e97c-4fbd-999e-cd951a478da2 · inbound

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments cites this paper.

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments Pix2seq: A Language Modeling Framework for Object Detection

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:17:23.693746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T21:15:38.699918Z digest=sha256:7e8d9165ba9aecb9096227d2190d1b6f3b5c3ded35a06c3396882796ddf76808

Observation 417e40cd-1d1d-48b8-8569-3011bd7c57b8 · inbound

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments cites this paper.

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments Pix2seq: A Language Modeling Framework for Object Detection

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:59:02.648707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T22:54:21.356538Z digest=sha256:bf66006e13e05c1e3eba4a08dedadf68bbc07dc6e07aa8ea8886cd482163c7e8

Observation 694233de-d16d-4ad9-bde4-cfad591daac5 · inbound

DeepGaze3.5-VL: Modeling Scanpaths via Autoregressive Token Prediction cites this paper.

DeepGaze3.5-VL: Modeling Scanpaths via Autoregressive Token Prediction Pix2seq: A Language Modeling Framework for Object Detection

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:58:37.580647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T15:49:17.217482Z digest=sha256:5d99d055876d23e777f1319e71cb2d5f353db6fcc4ba3cdac6b96232557412e3

Observation 07a0372d-df47-44a6-ba5e-1cc3835d7c86 · inbound

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement cites this paper.

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Pix2seq: A Language Modeling Framework for Object Detection

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-05T10:23:59.183387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:23:59.183387Z digest=sha256:d3470b65f32e884865877b86a5fca658811c35a081ccf9a740ca31da9fb1738e