Pith. sign in

REVIEW 2 cited by

Vision Transformer-based Semantic Communications With Importance-Aware Quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06038 v1 pith:3LLC76Z2 submitted 2024-12-08 eess.SP cs.CVcs.ITmath.IT

classification eess.SPcs.CVcs.ITmath.IT
keywords semanticimagecommunicationsquantizationallocationcommunicationframeworkproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic communications provide significant performance gains over traditional communications by transmitting task-relevant semantic features through wireless channels. However, most existing studies rely on end-to-end (E2E) training of neural-type encoders and decoders to ensure effective transmission of these semantic features. To enable semantic communications without relying on E2E training, this paper presents a vision transformer (ViT)-based semantic communication system with importance-aware quantization (IAQ) for wireless image transmission. The core idea of the presented system is to leverage the attention scores of a pretrained ViT model to quantify the importance levels of image patches. Based on this idea, our IAQ framework assigns different quantization bits to image patches based on their importance levels. This is achieved by formulating a weighted quantization error minimization problem, where the weight is set to be an increasing function of the attention score. Then, an optimal incremental allocation method and a low-complexity water-filling method are devised to solve the formulated problem. Our framework is further extended for realistic digital communication systems by modifying the bit allocation problem and the corresponding allocation methods based on an equivalent binary symmetric channel (BSC) model. Simulations on single-view and multi-view image classification tasks show that our IAQ framework outperforms conventional image compression methods in both error-free and realistic communication scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Importance-Aware Semantic Communication in MIMO-OFDM Systems Using Vision Transformer

    eess.SP 2025-08 unverdicted novelty 5.0 of 10

    A pretrained ViT's attention scores steer quantization, subcarrier mapping, and power allocation in MIMO-OFDM to improve semantic communication performance.

  2. Enhancing Wireless Networks for IoT with Large Vision Models: Foundations and Applications

    cs.NI 2025-08 conditional novelty 4.0 of 10

    The paper surveys LVM applications in wireless and reports a case study where progressive fine-tuning of pretrained LVMs gives more robust joint beamforming and positioning than a from-scratch CNN.

Pith tools