Pith. sign in

Paper Citation Record · LEDGER

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

As of 16 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2511.01390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.01390 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:26:34.931761Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13e18712-028d-4062-86a6-8106ca9f484a · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment UNITER: UNiversal Image-TExt Representation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.881258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.881258Z digest=sha256:6b97a40d4521513fc092916f3a700245b1a1e60e908412a6cf43b9849817afb7

Observation e8e8c1f8-b211-4a09-9b66-1dcaaccc404e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.890073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.890073Z digest=sha256:0d251b3e6bc8f3a123804495b81f4db589b2dddce4489783c2fa7a455800870e

Observation ef988689-8f4a-486f-b352-e90a489da990 · outbound

This paper cites Decoupled Weight Decay Regularization.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Decoupled Weight Decay Regularization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.907274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.907274Z digest=sha256:3336db9318c2ba1993d490310c95fe0408fd5ef643f94e6ea64261b32f4d7fe2

Observation 9ba09913-986d-473a-bb3a-c946c2025e09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Learning transferable visual models from natural language supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.915761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.915761Z digest=sha256:7d84b5e79ecab99d52e971395503ffd690f1c31a5346aac91b8d1a7fe439065c

Observation e41aaab0-26be-4162-815c-c83f2c3eb5ed · outbound

This paper cites Table 4: The comparisons of image-text retrieval for SEPS-Vit and SEPS-Swin with different selec- tion ratioρon Flicker30K.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Table 4: The comparisons of image-text retrieval for SEPS-Vit and SEPS-Swin with different selec- tion ratioρon Flicker30K

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.927587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.927587Z digest=sha256:c20dfaf36d2253a00346f016b5db8101edb374bc2371518df1b88ef26da9b58b

Observation 4e09f437-9490-4a72-a487-2cda4586577e · outbound

This paper cites an unresolved cited work.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.931761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.931761Z digest=sha256:8764d0ffd47b311a0af160cb2894319b59fc6343bcff0e686d1bb35293c8e0dc

Observation 56fa954e-cab7-4085-88b5-3f7fa5188f2f · outbound

This paper cites Image-question-answer synergistic network for visual dialog.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Image-question-answer synergistic network for visual dialog

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.898696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.898696Z digest=sha256:ec9f688c26fec255a809ce04f66f5d7a653074a78f9fd6a97af3ab77dda9223d

Observation 627a55b4-1432-4506-bed7-97b5580d9b60 · outbound

This paper cites The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.911429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.911429Z digest=sha256:3f5a4963a04e788bd27eb6d003e838af0bdd03e48237154f13b457f5a7e0b427

Observation 14faa27c-3a61-4dfa-9a5c-93b367c8ba76 · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.894136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.894136Z digest=sha256:184fe81a26affdca74961f03b58a43c6a9c3d0b0c9e1a4033755b64b390eaea2

Observation b6643d60-1288-4c2e-b079-7bd486bcfd1f · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.886208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.886208Z digest=sha256:03e99a3e32bd02273fbe9f675c3774896b19bff44a191e6f383b72336dff18e7

Observation 4fa0c16c-36bd-49c5-9b8b-91f1bd5ba792 · outbound

This paper cites SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.919618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.919618Z digest=sha256:4fbf8933d4e630448f45951606a16e537e3d37a25ac79c3797e211a25582ed24

Observation 2cd20453-59c6-4e10-af28-860e8b47f367 · outbound

This paper cites Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.902867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.902867Z digest=sha256:7b23927f82b5ec61990b995b4f1a45ae65d996353b0c32cfbd5e9e397c2122b6

Observation 1efe51f5-7dbe-4d7b-9141-9e9ba65ec5e8 · outbound

This paper cites In stark contrast, SEPS not only significantly surpasses all prior fine-grained methods but also successfully bridges this performance gap.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment In stark contrast, SEPS not only significantly surpasses all prior fine-grained methods but also successfully bridges this performance gap

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.923686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.923686Z digest=sha256:c65480fef0beba71fda0a39f449117d200fa638a1ddfcdfc7065dad8c0ea0aae

Pith citing papers

No inbound Pith citation observations are available.