Pith. sign in

Paper Citation Record · LEDGER

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers

As of 23 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 1 inbound Pith citation observation for arXiv:2501.09221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09221 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:12:58.172261Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:35:03.349259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T17:35:03.450625Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1966fa80-2ed0-491e-949d-13af8a6fcf86 · outbound

This paper cites Describing objects by their attributes.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Describing objects by their attributes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.355315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:12:58.140847Z digest=sha256:784ae7b1d3102ca95565f0946ceda6774513839dadc6045cca363a3065e06dc2

Observation 2fc6978c-abd0-4474-91ba-1071b70ea112 · outbound

This paper cites AST: Audio Spectrogram Transformer.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers AST: Audio Spectrogram Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.144966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.144966Z digest=sha256:bbc7feaac4a390e5f1aa4ce9486cbd076ebe0f2f1189a1a87e507df126641112

Observation 2af4640b-0220-4526-a915-86b3be921b4e · outbound

This paper cites Are Convolutional Neural Networks or Transformers more like human vision?.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Are Convolutional Neural Networks or Transformers more like human vision?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.153980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.153980Z digest=sha256:c3cf15b1f122cbea3b0532dbb35c1a6717ea18beb2d0370c38dc3bb178798fc8

Observation 1e30bcfb-b02e-4203-815c-e7bb59b0959e · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.163307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.163307Z digest=sha256:c39cbb2a9dedc83705ba6be1769afd828e03674a47620b945cb8984ad2eb718d

Observation 1e9718b5-4a7d-4876-a0d8-8e6e79d6de7c · outbound

This paper cites We set the patch size to correspond to 16x16 pixels.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers We set the patch size to correspond to 16x16 pixels

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.328782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:12:58.168030Z digest=sha256:f7bc25b0d4ec8fb343649f73f1825b2cd2fb2864ed35bd11241b93ea164c53fd

Observation 5aedb210-1e7d-4602-9b5f-bfde6765f8ba · outbound

This paper cites The following convolution blocks consist of a single convolution layer of kernel size=3 and stride=2, as well as the batch norm and ReLU activations.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers The following convolution blocks consist of a single convolution layer of kernel size=3 and stride=2, as well as the batch norm and ReLU activations

Reference 1029

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.316300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:12:58.172261Z digest=sha256:6096aa987f8b4feafd41394ff7abfa2411f10ea7a74438bca1de493e60703769

Observation 55785ad0-265a-4d90-a10f-99f7d97a28a6 · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers The caltech-ucsd birds-200-2011 dataset

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.341517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:12:58.158392Z digest=sha256:dcfa65de9d9c84997e308d8b5ec144e079f107f507f2d293b7f135a98fa967d1

Observation febaf998-ae44-47df-9776-593b65b844ff · outbound

This paper cites Image transformer for explainable au- tonomous driving system.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Image transformer for explainable au- tonomous driving system

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.370876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:12:58.136626Z digest=sha256:f5049fb744fd9b2c41380c647dc6aebb2601e4109aa8f92c385a3b3df0919730

Observation 654f5a22-82c6-4486-a61c-88f8aa928c89 · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Vision Transformer Adapter for Dense Predictions

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.127534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.127534Z digest=sha256:19cf7e2fac4303a2b1ca64bca5f789699d8fb88309ac809bbf990e4f93143dae

Observation a742862d-fc28-4129-aaf0-71e73f46c555 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.131823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.131823Z digest=sha256:8a55118f94c3e80313c6170442f486cbc53971d972ff32c150d6dc88073fef41

Observation d54e9a35-8411-4432-82d3-9ea3956193b6 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers On the Opportunities and Risks of Foundation Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.123392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.123392Z digest=sha256:8e5672ec0cf7ce1d54f4652b97acf49573bdaf5c74c53efb2c62112544281e5b

Observation ec637453-89cc-459b-88f8-9e10cb04b1c7 · outbound

This paper cites SiT: Self-supervised vIsion Transformer.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers SiT: Self-supervised vIsion Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.118649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.118649Z digest=sha256:2b36ae4b390290457e254907ad067363f088793ddb3b820233a5936c7072662a

Observation df43ab0e-c0d5-4cb3-aa0a-487c8a98c86d · outbound

This paper cites A Self-explaining Neural Architecture for Generalizable Concept Learning.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers A Self-explaining Neural Architecture for Generalizable Concept Learning

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:12:58.236527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:12:58.149323Z digest=sha256:d6131b1c95c352fddf33d079c5705029141c222d0a3feef6a1ccf4512d676aa9

Pith citing papers

Observation d069482b-01f8-4bcf-8cb7-08316db20626 · inbound

Integrating attention into explanation frameworks for language and vision transformers cites this paper.

Integrating attention into explanation frameworks for language and vision transformers ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:35:03.459622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:35:03.349259Z digest=sha256:38a8d47f64988bf9291941bf567ef2ac84b666b106a964c2ebd1b38f5c0fbf21