Pith. sign in

Paper Citation Record · LEDGER

TULIP: Towards Unified Language-Image Pretraining

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2503.15485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15485 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:57.321888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.381543Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4293610-9839-41e7-889f-03aa336b86c7 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence TULIP: Towards Unified Language-Image Pretraining

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.879831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:9b0b48e84131386e24a38e8e6c79815e3bc200e7685e6eaf3ea9e106e1785e21

Observation b78fa165-3361-4744-9e81-cf90a451b496 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence TULIP: Towards Unified Language-Image Pretraining

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.385092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:2a2ffb538f635ed9471db9bbee4dd690b9768c96a2a38cb5e10d7088690916a9

Observation 1ca84f57-c161-401b-939f-f32b0ecf4a29 · inbound

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets cites this paper.

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets TULIP: Towards Unified Language-Image Pretraining

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:57.321888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:57.321888Z digest=sha256:c15a28a81d0e1a994f1798030e75a902c87d4900029c44fe59c36e4cc48fd095

Observation 8ffa87e2-96fa-4345-8c8b-49c5e676b941 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning TULIP: Towards Unified Language-Image Pretraining

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.682149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:dcbe8a5eb6804ae5d10cd6f5fa9cca732318376172420a54801eb8abd393175c

Observation 0519322e-3314-4085-ba16-d45541650657 · inbound

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics cites this paper.

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics TULIP: Towards Unified Language-Image Pretraining

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:40.606092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T04:25:58.275705Z digest=sha256:48d5e44a292be8400d54f59a416d4abdbf06e3ba20a909cb2c0f8de9a20c72b9

Observation 13e07a65-2d5a-469d-bf35-4e99ee468c0e · inbound

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models cites this paper.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models TULIP: Towards Unified Language-Image Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.402513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.402513Z digest=sha256:a4c073a8dd1a799b94147ff714c8b726559f0700fa99c714194772ddb234e85b