Pith. sign in

Paper Citation Record · LEDGER

TULIP: Towards Unified Language-Image Pretraining

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2503.15485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15485 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:57.321888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.381543Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4293610-9839-41e7-889f-03aa336b86c7 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence TULIP: Towards Unified Language-Image Pretraining

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.879831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:0ab7e25110feb84bffbdabc3784b886bf4f743ff90aef7af7f0b41fb453bff5e

Observation b78fa165-3361-4744-9e81-cf90a451b496 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence TULIP: Towards Unified Language-Image Pretraining

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.385092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:148ef80ea1b7166b7e06f68fbeb1bf50e5505f2c3004762303b0080bdda2cda9

Observation 1ca84f57-c161-401b-939f-f32b0ecf4a29 · inbound

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets cites this paper.

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets TULIP: Towards Unified Language-Image Pretraining

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:57.321888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:57.321888Z digest=sha256:be601e301421403c49548cc810ddaca7a2f6908f1b1a2cb5903a6284fad3d40d

Observation 8ffa87e2-96fa-4345-8c8b-49c5e676b941 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning TULIP: Towards Unified Language-Image Pretraining

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.682149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:9f09204006fc4c111fb37f3e8c1e218f4059f7649e456c4dc68a605d246690f9

Observation 0519322e-3314-4085-ba16-d45541650657 · inbound

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics cites this paper.

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics TULIP: Towards Unified Language-Image Pretraining

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:40.606092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T04:25:58.275705Z digest=sha256:4cfd6dd67e3a41e140ff342a9140c01095feffaa8c8d18744600dd591d8a87b3

Observation 13e07a65-2d5a-469d-bf35-4e99ee468c0e · inbound

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models cites this paper.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models TULIP: Towards Unified Language-Image Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.402513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.402513Z digest=sha256:8de1caf566a4de8186b8730d75736069c364d5b495adb910f2f243052465a138