REVIEW 19 cited by
End-to-end Optimized Image Compression
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We describe an image compression method, consisting of a nonlinear analysis transformation, a uniform quantizer, and a nonlinear synthesis transformation. The transforms are constructed in three successive stages of convolutional linear filters and nonlinear activation functions. Unlike most convolutional neural networks, the joint nonlinearity is chosen to implement a form of local gain control, inspired by those used to model biological neurons. Using a variant of stochastic gradient descent, we jointly optimize the entire model for rate-distortion performance over a database of training images, introducing a continuous proxy for the discontinuous loss function arising from the quantizer. Under certain conditions, the relaxed loss function may be interpreted as the log likelihood of a generative model, as implemented by a variational autoencoder. Unlike these models, however, the compression model must operate at any given point along the rate-distortion curve, as specified by a trade-off parameter. Across an independent set of test images, we find that the optimized method generally exhibits better rate-distortion performance than the standard JPEG and JPEG 2000 compression methods. More importantly, we observe a dramatic improvement in visual quality for all images at all bit rates, which is supported by objective quality estimates using MS-SSIM.
Forward citations
Cited by 19 Pith papers
-
End-to-end image compression and reconstruction with ultrahigh speed and ultralow energy enabled by opto-electronic computing processor
An integrated optoelectronic processor with a programmable 32x32 photonic matrix performs end-to-end image compression and reconstruction at 49.5 ps/pixel and 10.58 nJ/pixel, according to the authors.
-
Compress-Align-Detect: onboard change detection from unregistered images
A single neural network performs compression, co-registration, and change detection onboard a satellite, achieving F1 up to about 70% at low bitrates on simulated unregistered image pairs.
-
Locality-Aware Density Control for Efficient Gaussian-based Image Representation
A locality-aware density-control framework for 2D Gaussian image representation that densifies coherent high-error regions and merges redundant similar Gaussians, improving PSNR at fixed budgets.
-
-8 dB SNR + 90% Packet Loss: MamVSC -- CSI-Guided Semantic Mamba for Extreme-Robust Video Semantic Communication
A Mamba-based semantic video communication system with CSI-guided adaptive encoding and packet loss recovery achieves PSNR > 21 dB at -8 dB SNR and 90% packet loss in AWGN channels.
-
HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
A hyperprior predicts a Gaussian in codebook space and converts it to index probabilities, enabling content-adaptive entropy coding for VQ image compression.
-
SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
A semantic communication framework uses MLLM-derived scenario-aware importance labels to allocate coding resources, improving PSNR of important image regions at comparable or lower bandwidth than prior JSCC systems.
-
JPEG Processing Neural Operator for Backward-Compatible Coding
JPNeO improves JPEG compression with neural operators at encoding and decoding, without changing the JPEG bitstream format.
-
LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression
A lossless point cloud geometry codec overfits a small sparse-convolution network per group of frames and entropy-codes octree occupancy using the network's predicted probabilities.
-
CosmoFlow: Scale-Aware Representation Learning for Cosmology with Flow Matching
Flow matching on CDM fields yields an 8 number latent that reconstructs fields and estimates Omega_m and sigma_8 nearly as well as a raw-field network, with channels tied to spatial scales.
-
StableCodec: Taming One-Step Diffusion for Extreme Image Compression
A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.
-
End-to-End RGB-IR Joint Image Compression With Channel-wise Cross-modality Entropy Model
A channel-wise cross-modality entropy model with low-frequency context fusion improves joint RGB-IR image compression, achieving 23.1% bit rate savings over the previous state of the art on LLVIP.
-
Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution
A single rate-variable generative compression model treats quantization as a forward corruption and reverses it with a two-step denoiser, outperforming prior generative codecs on perceptual quality benchmarks.
-
Neural Video Compression with Context Modulation
DCMVC modulates the propagated temporal context with an additional oriented context from the reference frame, reporting 10.1 percent bitrate savings over DCVC-FM and 22.7 percent over VVC on standard test sets.
-
SDGIC: A Semantic Disambiguation-Guided Generative Image Compression Method for Ultra-Low Bitrates
A diffusion-based image codec guided by text, a highly compressed image, and CLIP-derived semantic pseudo-words improves semantic consistency at bitrates below 0.05 bpp.
-
RAVQ-HoloNet: Rate-Adaptive Vector-Quantized Hologram Compression
RAVQ-HoloNet compresses phase-only holograms with a hierarchical VQ-VAE whose codebook size is shrunk by an LSTM module, enabling multiple bitrates from one model and claiming BD-Rate -33.91% vs DPRC.
-
Efficient Learned Image Compression Through Knowledge Distillation
Knowledge-distilled students with 64 or more channels match the rate-distortion performance of a 128-channel teacher while cutting memory by 68% and energy by 34%.
-
Conquering High Packet-Loss Erasure: MoE Swin Transformer-Based Video Semantic Communication
MSTVSC is a packet-loss-resistant semantic video codec that recovers erased semantic elements with a 3D CNN and claims MS-SSIM above 0.6 and PSNR above 20 dB at 90% packet loss.
-
Explicit Residual-Based Scalable Image Coding for Humans and Machines
Explicitly compressing pixel or feature residuals between machine-oriented and human-oriented image codec layers improves scalable coding efficiency, with up to 29.57% BD-rate savings over ICMH-FF.
-
Point Cloud Compression and Objective Quality Assessment: A Survey
A survey of point cloud compression and objective quality assessment that benchmarks representative methods on standard datasets and distills design insights.
Discussion (0). Continue with ORCID to comment.