Pith. sign in

REVIEW 2 cited by

Hyper-Representations for Pre-Training and Transfer Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.10951 v1 pith:P4OCF7WM submitted 2022-07-22 cs.LG

classification cs.LG
keywords modelhyper-representationsmodelslearningknowledgeneuralpotentialpre-training
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Learning representations of neural network weights given a model zoo is an emerging and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we extend hyper-representations for generative use to sample new model weights as pre-training. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and a sampling method based on the empirical density of hyper-representations. The models generated using our methods are diverse, performant and capable to outperform conventional baselines for transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A prompt-conditioned hyper-network generates LoRA fine-tuning weights for unseen tasks in a single forward pass, without training on the target dataset.

  2. Recurrent Diffusion for Large-Scale Parameter Generation

    cs.LG 2025-01 conditional novelty 6.0 of 10

    RPG generates full weights for models up to 200M parameters, including ConvNeXt-L and LLaMA LoRA adapters, at accuracy comparable to trained checkpoints, using recurrent token prototypes to condition a 1D diffusion model.

Pith tools