Pith. sign in

REVIEW 2 cited by

MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.02427 v2 pith:3SXKDZ3E submitted 2025-01-05 cs.CV

classification cs.CV
keywords videovideosmetanervneuralguidancerepresentationrepresentationsframes
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Neural Representations for Videos (NeRV) has emerged as a promising implicit neural representation (INR) approach for video analysis, which represents videos as neural networks with frame indexes as inputs. However, NeRV-based methods are time-consuming when adapting to a large number of diverse videos, as each video requires a separate NeRV model to be trained from scratch. In addition, NeRV-based methods spatially require generating a high-dimension signal (i.e., an entire image) from the input of a low-dimension timestamp, and a video typically consists of tens of frames temporally that have a minor change between adjacent frames. To improve the efficiency of video representation, we propose Meta Neural Representations for Videos, named MetaNeRV, a novel framework for fast NeRV representation for unseen videos. MetaNeRV leverages a meta-learning framework to learn an optimal parameter initialization, which serves as a good starting point for adapting to new videos. To address the unique spatial and temporal characteristics of video modality, we further introduce spatial-temporal guidance to improve the representation capabilities of MetaNeRV. Specifically, the spatial guidance with a multi-resolution loss aims to capture the information from different resolution stages, and the temporal guidance with an effective progressive learning strategy could gradually refine the number of fitted frames during the meta-learning process. Extensive experiments conducted on multiple datasets demonstrate the superiority of MetaNeRV for video representations and video compression.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MSNeRV is an implicit neural representation video codec that combines temporal-window fusion, GoP-level background grids, multi-resolution supervision, and multi-scale feature blocks, reporting strong compression resu...

  2. ImputeINR: Time Series Imputation via Implicit Neural Representations for Disease Diagnosis with Missing Data

    cs.LG 2025-05 conditional novelty 5.0 of 10

    ImputeINR uses implicit neural representations to impute missing time series values, reporting better MSE and MAE than nine baselines on eight datasets, with the largest gains at high missing rates.

Pith tools