Pith. sign in

REVIEW 1 cited by

Energy-Based Models for Code Generation under Compilability Constraints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.04985 v1 pith:UMIYPX4T submitted 2021-06-09 cs.LG cs.CLcs.NEcs.SE

classification cs.LGcs.CLcs.NEcs.SE
keywords codecompilabilitymodelcompilableconstraintenergy-basedgenerativemodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Neural language models can be successfully trained on source code, leading to applications such as code completion. However, their versatile autoregressive self-supervision objective overlooks important global sequence-level features that are present in the data such as syntactic correctness or compilability. In this work, we pose the problem of learning to generate compilable code as constraint satisfaction. We define an Energy-Based Model (EBM) representing a pre-trained generative model with an imposed constraint of generating only compilable sequences. We then use the KL-Adaptive Distributional Policy Gradient algorithm (Khalifa et al., 2021) to train a generative model approximating the EBM. We conduct experiments showing that our proposed approach is able to improve compilability rates without sacrificing diversity and complexity of the generated samples.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Incorporating Inductive Biases to Energy-based Generative Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Augmenting a neural energy-based model with a linear statistic term that encodes known data properties improves generation quality on molecules, digits, and point clouds.

Pith tools