Pith. sign in

Paper Citation Record · LEDGER

Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.17510.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.17510 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:00:53.613167Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T21:06:50.891380Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2ad3b221-e083-4d99-8ef1-4e2d65c5a961 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.613167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.613167Z digest=sha256:304ae190a840831d081d22ba3a29fd370efcaf0db8967cef4570c732dec49126

Observation a0f187c4-8a2f-4b11-acc3-86415fc8e676 · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.894430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:e9b455c0a662c5152342938307dd367b8be2eda7bcbfe453a2771b5c706cbee3

Observation 38cd6f60-4da3-4ec5-b0d7-c3d77a9fb7d5 · inbound

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance cites this paper.

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T05:16:57.528659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:16:57.528659Z digest=sha256:3666024e09ab5a2dbce6d86a67c92bda8e4ae1c3fcb016f1396723a5bc82c106

Observation f8928050-f547-4afd-bb8f-4c6f38367534 · inbound

Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization cites this paper.

Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:12:54.490142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:12:54.490142Z digest=sha256:1e86b60ae3584f1c40ec6ab2951f9787a61c16a0f81a56861882292e3536a279

Observation 6c5f371e-347c-4a33-8e2b-e8d1eb724d91 · inbound

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States cites this paper.

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T16:52:57.036688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:52:57.036688Z digest=sha256:fdd41bf78f0c3df151acda6893d69028c92dabdaa9aa672c47b4310649227a26

Observation 83e40a77-496d-47c9-9b1d-76c8fb930a11 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:46.820969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:b7d82b405a34deac96618cfe96d7838acac4d637886d764b33b36027dbb6b2d6

Observation 52b56ada-bc48-42b5-8437-6b8dd70795af · inbound

Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding cites this paper.

Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:51.851266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:51.851266Z digest=sha256:1769292ca2e5877395dfacc42ad60dfc6a2bedbe98a823dea789620fbb3647ab