Pith. sign in

Paper Citation Record · LEDGER

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2511.01390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.01390 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:26:34.931761Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13e18712-028d-4062-86a6-8106ca9f484a · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment UNITER: UNiversal Image-TExt Representation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.881258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.881258Z digest=sha256:3a4ad5e2dbceb2d487ef7ecc9df11484007734883a59a56dd5ea89d5b5824328

Observation e8e8c1f8-b211-4a09-9b66-1dcaaccc404e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.890073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.890073Z digest=sha256:7aa5271ddc7f8d023e08d04655139890b5ecd02da8b6c8aae9a8494246a3b6a4

Observation ef988689-8f4a-486f-b352-e90a489da990 · outbound

This paper cites Decoupled Weight Decay Regularization.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Decoupled Weight Decay Regularization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.907274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.907274Z digest=sha256:51c28809c164b734bf82c52cfbb61664ec09e2801cc6175605756cf51258269b

Observation 9ba09913-986d-473a-bb3a-c946c2025e09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Learning transferable visual models from natural language supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.915761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.915761Z digest=sha256:66d6e1c1a972a0f35d0a73e0b7cab44be4c9d9b5593bbfaf9d701f9185f1a54b

Observation e41aaab0-26be-4162-815c-c83f2c3eb5ed · outbound

This paper cites Table 4: The comparisons of image-text retrieval for SEPS-Vit and SEPS-Swin with different selec- tion ratioρon Flicker30K.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Table 4: The comparisons of image-text retrieval for SEPS-Vit and SEPS-Swin with different selec- tion ratioρon Flicker30K

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.927587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.927587Z digest=sha256:24c18988a8c07dedf5d27b997d9825a1d3b5dad927ac7e1e9ad7c25526ccabec

Observation 4e09f437-9490-4a72-a487-2cda4586577e · outbound

This paper cites an unresolved cited work.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.931761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.931761Z digest=sha256:760a7b5ece40958b61c11d592921ec640f6c50bc11ab53c4c2619ac07e0ef71b

Observation 56fa954e-cab7-4085-88b5-3f7fa5188f2f · outbound

This paper cites Image-question-answer synergistic network for visual dialog.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Image-question-answer synergistic network for visual dialog

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.898696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.898696Z digest=sha256:1cf784446f7f24d09edd72ab557557359955d8e4231f57f0d2f42efa33dfc2fb

Observation 627a55b4-1432-4506-bed7-97b5580d9b60 · outbound

This paper cites The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.911429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.911429Z digest=sha256:e6ca4ab6dbce725caef617e3e85f5429ab07ac55b178c631fd45161d1cd90a7c

Observation 14faa27c-3a61-4dfa-9a5c-93b367c8ba76 · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.894136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.894136Z digest=sha256:e1016cf6aeb29012c49a6eeef9999c5705707b1efb716d758ccf29f81469c982

Observation b6643d60-1288-4c2e-b079-7bd486bcfd1f · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.886208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.886208Z digest=sha256:11c09e72c4c1bba45e7e1164a0382d8ed1ed5aa24cfffd098681efefa89bee8c

Observation 4fa0c16c-36bd-49c5-9b8b-91f1bd5ba792 · outbound

This paper cites SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.919618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.919618Z digest=sha256:0d73f9a64c9f4c5899d442242caef8c16cb29749ad1eabf2cd136c5305eac8a7

Observation 2cd20453-59c6-4e10-af28-860e8b47f367 · outbound

This paper cites Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.902867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.902867Z digest=sha256:8c5358938d6b80521ee44fe14f1ac3ec2670762bf97463dcf8e59d2b9874d3bb

Observation 1efe51f5-7dbe-4d7b-9141-9e9ba65ec5e8 · outbound

This paper cites In stark contrast, SEPS not only significantly surpasses all prior fine-grained methods but also successfully bridges this performance gap.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment In stark contrast, SEPS not only significantly surpasses all prior fine-grained methods but also successfully bridges this performance gap

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.923686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.923686Z digest=sha256:7f7bfa027823baf12f0a4ccb09cd483d6b970239b728162a657021a648b4de7b

Pith citing papers

No inbound Pith citation observations are available.