Pith. sign in

Paper Citation Record · LEDGER

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation

As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2502.02763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02763 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:19:23.191967Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:18:06.989944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T00:19:46.820702Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aef2c08a-e2ef-4762-bcaf-da5d24a04762 · outbound

This paper cites Focal sparse convolutional networks for 3d object detection.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Focal sparse convolutional networks for 3d object detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.532068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.079589Z digest=sha256:51cc2663d59d151885ce9231d752ad078f1bb453b712f214592f6e06da090232

Observation 6b9a0d16-b217-4c80-a1e9-e18b7178aeaf · outbound

This paper cites Deformable convolutional networks.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Deformable convolutional networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.516609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.084083Z digest=sha256:901eb3018be805ca2306d2cde0ae5ca7ab29b0550883019b1c4fe9a94a28b3e9

Observation e692c0da-2b99-498e-ab61-ddf0e784fd93 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Objaverse: A universe of annotated 3d objects

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.505803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.088072Z digest=sha256:09537751ae118ff383f6d9df2aadc6bd1d3698f3b45559a6660b8e70f1795e0f

Observation 5a218b3b-ef72-4de6-a4a2-0ea29ab5a7e2 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.091877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.091877Z digest=sha256:8ca1e5a0da841bd19a9e34841532f262eea01d86af5066d4f5fa170bcc5eec9e

Observation ed9c03ac-82bf-4c07-8931-eda19d4fd0ee · outbound

This paper cites C., and Kipf, T.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation C., and Kipf, T

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.494268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.096191Z digest=sha256:073da5a7da52c7fa2b6392aa7cde98a6f9fad39319933112e7ffe2cc47744d0b

Observation e355b5a9-61dc-4867-bed8-e3bc0e20e8dd · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Lvis: A dataset for large vocabulary instance segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.099713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.099713Z digest=sha256:ccc3d47f758c40bb9f2a3c2af4c39c8761565dfa00f5d27c6ad36a32ab1cba8c

Observation 63c4ade8-74ca-41fa-acb4-3271184e0ee0 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.103509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.103509Z digest=sha256:fd069d079b4c0a905ba45ee6b1f6d35f0afbd43a369f3973e768562e7e7af42c

Observation d6085e78-f0da-42b2-80e4-1bc5be682996 · outbound

This paper cites S., Sochenov, A., Leimk \"u hler, T., Okunev, M., Goodall, T., and Rufo, G.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation S., Sochenov, A., Leimk \"u hler, T., Okunev, M., Goodall, T., and Rufo, G

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.477171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.106998Z digest=sha256:312e8194de6423e4ccaf1571c9eae81f5ef57d4ddcfaf2e6e6fb41a33556ad71

Observation 348ee7a9-8c1f-46a1-baf1-ac6afdd9af4c · outbound

This paper cites C., Lo, W.-Y., et al.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation C., Lo, W.-Y., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.110318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.110318Z digest=sha256:0294191205a654062a2da76a56fedebb4d86741b96d2f255de8e45b36373b277

Observation 2727d70f-9aa7-45b2-9a47-13d1e7b60a63 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.113657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.113657Z digest=sha256:d2243026b75c7e84734cf5439651e91658a055c314d22df7b6f66a85d35b2b2f

Observation 61d094aa-1c0b-476c-a91f-1cb4287a185e · outbound

This paper cites Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.453297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.117269Z digest=sha256:f24d589db5a4ddbccfb0394e8e95fa260226f73d7c0fcb5f878b2968d5ab2f5f

Observation dfcf102c-27e3-47a9-b365-35c904a1c5fe · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.120823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.120823Z digest=sha256:0da14d1ccad658d19f786b0b74b0f8cf4f059cdcf70e46240698912c8919d329

Observation ecbb7336-e589-441a-8320-b7a771ff23a5 · outbound

This paper cites A convnet for the 2020s.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation A convnet for the 2020s

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.124894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.124894Z digest=sha256:1bf64d9f7ca1c4bf71a005db0268cbf719933d4fca82f1c51905bb62339e9fe8

Observation 4761056f-94e7-4bdf-86a2-2f0bac12a388 · outbound

This paper cites Object-centric learning with slot attention.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Object-centric learning with slot attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.128418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.128418Z digest=sha256:64272f860e5f01ff3c493a01582550e7e13d8d3bb0df881d686f61a15faf7f43

Observation fa0d062a-7b94-4f7a-95b8-f599a6995fef · outbound

This paper cites Biologically inspired deep learning model for efficient foveal-peripheral vision.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Biologically inspired deep learning model for efficient foveal-peripheral vision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.423820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.132289Z digest=sha256:4e8af0ff71e0e678ba4de972c314316a92e310c6b5b3b59c371963e215d996c4

Observation 39862b7b-c7af-4362-bf18-b8b4f76699e2 · outbound

This paper cites A., Paczan, N., Webb, R., and Susskind, J.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation A., Paczan, N., Webb, R., and Susskind, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.412652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.136215Z digest=sha256:ea9e8056d84a93a49a6c8abde64db0bf891b17182581779f63b29cd14f8e095f

Observation 3de9cd10-298a-4932-87b4-247187451080 · outbound

This paper cites Simple unsupervised object-centric learning for complex and naturalistic videos.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Simple unsupervised object-centric learning for complex and naturalistic videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.401449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.140120Z digest=sha256:eac16f15b2135c2b26fe611bbb0e6bf4d0b500bf5ab19e8f04b223e438e6067f

Observation f73b75ad-9853-491b-b4fb-24f6c98c6f8b · outbound

This paper cites Fovea: Foveated image magnification for autonomous navigation.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Fovea: Foveated image magnification for autonomous navigation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.389788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.143775Z digest=sha256:d0b9afa9865268ab704a6be533196c13689e49bf18292ec487c07b5ecc756eb9

Observation 5ebfc136-c11f-4d3a-8233-7baeeafc4e15 · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:19:23.378406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.147186Z digest=sha256:3c35e2968f7543ada0fb70a668f6a05b5739a4ee9f0bf51e4c6375568715dee2

Observation ad15cafb-baf6-44eb-a9fa-b5ec089c81c3 · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:19:23.366867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.150628Z digest=sha256:ce21bae83d872871cda3a7f029ed133a7ff9b403ee14cfe661e7bc918fecfc78

Observation 11f94f2c-7b34-40ac-829c-2a84283a96fc · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:19:23.355602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.154164Z digest=sha256:d93a28779c05246810e44beb717788ea319fe760700726de83a167d60f4a0889

Observation de4422ff-b813-4f9b-8bd6-1d03b448c22e · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Tinyvit: Fast pretraining distillation for small vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.344416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.157736Z digest=sha256:9877acd60c8b6c8d08d866903184d4a8a128abe21c57f6dd7ee89381b0d6c360

Observation b1807004-4610-4e87-9e7d-d8620f8f41f1 · outbound

This paper cites EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.161236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.161236Z digest=sha256:4f41ef743c63af932500b1647fb1167af331fcb326d1d17991d037108ba85b39

Observation 5022d81f-a49d-4d69-bf07-c58193a110cc · outbound

This paper cites Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.333517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.165209Z digest=sha256:d3a758219915700b4bc6bb81e004fb7d8daea6f9d48120c557781223eaf78dc2

Observation 1283680c-c063-4688-ac1a-a4cec288a2b4 · outbound

This paper cites Lape: Layer-adaptive position embedding for vision transformers with independent layer normalization.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Lape: Layer-adaptive position embedding for vision transformers with independent layer normalization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.321516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.168897Z digest=sha256:9e050af1bea31fd94f3dea6c81742194a1023daddbd7ad70c63791ef3b08e571

Observation 572d0087-8a8b-49e8-9367-5a618df5a46b · outbound

This paper cites Hdri haven, 2016.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Hdri haven, 2016

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.309605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.172336Z digest=sha256:908ca833499a0859855aeb370178ac45662a90a307701c4bf2a1aa14c3f8853b

Observation 3526f2cc-6731-4960-a531-8761ca17a5f1 · outbound

This paper cites Object-centric learning for real-world videos by predicting temporal feature similarities.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Object-centric learning for real-world videos by predicting temporal feature similarities

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.297696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.176069Z digest=sha256:987457444b4dcf63a2d7b4a7fc9bce552d533b7a78bc1147b8da1d5e784468a6

Observation 456ce933-29a1-431c-87ba-1339b93fbc60 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.179663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.179663Z digest=sha256:1f37029a850db6d1fe410ae2b0f3a11ca4e7bdaaf1e10d48e69f01c39489e51b

Observation ecee854f-05e6-4dcc-8e8e-7774c690e1d5 · outbound

This paper cites Fast segment anything, 2023.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Fast segment anything, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.286527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.183512Z digest=sha256:2f604b8a5e2e15fd603f2934c26583b8003f3dd74c9e70ca3141cbca93236119

Observation 698f2ba0-d7ed-4fcd-9f6d-e244b795f22b · outbound

This paper cites Deformable convnets v2: More deformable, better results.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Deformable convnets v2: More deformable, better results

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.275476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.188345Z digest=sha256:57fea19dba2d618d508298b248672f1ab30f051e706ec0642292bc60ac23f014

Observation a8a9c764-eca2-407d-87bc-0a74a9bf7549 · outbound

This paper cites write newline.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation write newline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.191967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.191967Z digest=sha256:d52009695e6cf2521484a11a44441a688fddb851d577def8df9dfc7b02cf4e27

Pith citing papers

Observation be67138d-9cdd-4c25-973e-ac8abc8f5a1d · inbound

Self-supervised pretraining for an iterative image size agnostic vision transformer cites this paper.

Self-supervised pretraining for an iterative image size agnostic vision transformer Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-13T01:17:40.500289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:18:06.989944Z digest=sha256:e3fe5ac6d0203882d641054ea53018665c84290a25edcea2ef06a939d665f8b3