ACL-2020 Contrastive Code Representation Learning
ContraCode: MoCo-based contrastive learning for code
Abstract
The authors observe that RoBERTa is overly sensitive to edits of source code even when those edits preserve semantics. They propose ContraCode, a pretraining method that learns to recognize semantically similar variants.
Introduction

Under semantics-preserving adversarial modifications, RoBERTa performs worse than random classification.
The authors apply source-to-source, compiler-based transformation techniques—for example, deleting “dead” code by removing operations that do not change the program’s output.
Model
Code transformation techniques fall into three categories:
- Code minimization: alter syntactic structure and apply corrective constructive transforms, such as precomputing constant expressions
- Identifier renaming: randomly substitute method and variable names
- Normalization transforms: improve generalization by reducing trivial positive pairs with high textual overlap
A transformation discard algorithm enforces diversity, mainly to ensure that transformed code is actually modified. After 20 random sequential transformations, the authors find that 89% of methods yield more than one distinct alternative.

Contrastive pretraining:
- Extends MoCo by forming positive pairs from the same program and negative pairs from different programs
- Uses a queue to store negative samples, as in MoCo
- A large negative set of more than 100k examples
ContraCode is agnostic to the encoder architecture; the authors evaluate a two-layer bidirectional LSTM and a six-layer Transformer.

Evaluation
Experiments cover three settings: zero-shot clone detection, fine-tuned type inference, and extreme code summarization.
zero-shot Code Clone Detection


Fine-tuning for Type Inference


Extreme Code Summarization


Understanding augmentation importance
For sequence-to-sequence summarization, the following augmentation techniques are used:
- Line subsampling (LS): randomly sample a subset of lines from a method (P=0.9); this does not preserve semantics and acts as a regularizer
- Subword regularization (SW): tokenize text into varying forms—whole words or subtokens
- Variable renaming (VR), identifier mangling (IM): rename variables and identifiers
- Dead-code insertion (DCI): insert no-op or log statements
For type inference, LS and SW are used.
Additional code augmentation techniques:
- Dead-code elimination (DCE): remove useless code
- Type upconversion (T): e.g., replace
&withandortruewith1 - Constant folding (CF): precompute expressions, e.g., replace
(2+3) * 4with20
<HR align=left color=#987cb9 SIZE=1>