Pith. sign in

Paper Citation Record · LEDGER

Improved baselines for vision-language pre-training

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2305.08675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.08675 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:22.515438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T13:58:16.333412Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bf53a0a5-6d35-4b63-9e11-ce2d0c64349b · inbound

Multimodal Autoregressive Pre-training of Large Vision Encoders cites this paper.

Multimodal Autoregressive Pre-training of Large Vision Encoders Improved baselines for vision-language pre-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.515438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.515438Z digest=sha256:c18ae4487f08a44acc7b661a5106198ccbad5c7d5dfc248683083be659b71a2c

Observation 9d6152d3-944f-4766-ba2c-87ff0c6d2b23 · inbound

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference cites this paper.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Improved baselines for vision-language pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.887371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.887371Z digest=sha256:dc908a8225332fefa40b09fd8c910db57a60265b64161c445001a803da4ecc87

Observation 048aaf41-1698-41c2-aa7e-09ab37edfb3e · inbound

HyperCLIP: Adapting Vision-Language models with Hypernetworks cites this paper.

HyperCLIP: Adapting Vision-Language models with Hypernetworks Improved baselines for vision-language pre-training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T10:19:24.021583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:19:24.021583Z digest=sha256:b6af373464356c6b34c8d8abce66c507f339c3c717e737ef7f40942a30fa0737

Observation 69cdd656-8c2b-477e-a295-e79ba726b29a · inbound

TAPS : Frustratingly Simple Test Time Active Learning for VLMs cites this paper.

TAPS : Frustratingly Simple Test Time Active Learning for VLMs Improved baselines for vision-language pre-training

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:58:16.459417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:58:11.499233Z digest=sha256:23a23e7db35f1fdb96207da29cf5afbeb701c0b9d5eaa1661ba9d891a5da2992