IJCNLP-2017 A parallel corpus of Python functions and documentation strings for automated code documentation and code generation
code-docstring-corpus: a Python code–docstring generation dataset
Abstract
The authors present a large, diverse parallel corpus of Python functions and documentation strings, and experiment with data augmentation methods to increase the volume of training data.
Introduction
Existing corpora
Some existing corpora—such as the DJANGO dataset and the Project Euler dataset—are human annotated and yield high-precision examples, but at the cost of expensive annotation and small (from hundreds to fewer than 20,000 examples), homogeneous datasets.
Other corpora match user-generated descriptions with code snippets mined from public websites such as StackOverflow or IFTTT. These datasets can be very large (>100k) but are typically highly noisy.
Still other resources are tailored to specific domains.
Each of these corpora has limitations. For example, the DJANGO and Project Euler corpora use pseudocode rather than genuine natural language as code descriptions, which makes code snippets and descriptions structurally similar and easy to align. In the Magic the Gathering and Hearthstone corpora, code snippets are duplicated, and most examples are either boilerplate or align with specific keywords in the description in only a limited way.
Our proposal
The authors therefore propose a code corpus in which each example includes a docstring; they extract only top-level functions and attach metadata (repository owner, repository name, etc.) so that contextual relationship graphs can be reconstructed.
Dataset



Baseline results

<HR align=left color=#987cb9 SIZE=1>