Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, Chengxu Zhuang

Motivation

Huge efforts have been put into optimizing LM pretraining at massive scales in the last several years (Raffel et al., 2020; Brown et al., 2020; Chowdhery et al., 2022; Hoffmann et al., 2022). While growing parameter counts often get the most attention, datasets have also grown by orders of magnitude. These increasingly larger pretraining datasets are visualized, to scale, in Figure 1. At the same time, there has been almost no progress in pretraining at smaller human-like data scales.

Focusing on scaled-down pretraining has several potential benefits: First, small-scale pretraining can be a sandbox for developing novel techniques that improve data efficiency. These techniques have the potential to then scale up to larger datasets commonly seen in applied NLP, and could be used to enhance current approaches to modeling low-resource languages. Second, improving our ability to train LMs on the same types and quantities of data that humans learn from will give us greater access to more plausible cognitive models of humans and help us understand what allows humans to acquire language so efficiently (Keller, 2010; Dupoux, 2018). That is, even model failure can help in developing hypotheses about the differences between human and LM language learning.

The goal of this shared task will be to incentivize researchers with an interest in pretraining and/or cognitive modeling to focus their efforts on optimizing pretraining given data limitations inspired by human development. Additionally, we hope to democratize research on pretraining—which is typically thought to be practical only for large industry groups—by drawing attention to open problems that can be addressed on a university budget.

Key Dates

HTML]E9F0E9 • January 2023: Training data released • March 2023: Evaluation pipeline released • July 15, 2023: Results due • August 1, 2023: Paper submissions due • Date TBA: Presentation at CoNLL

Tracks

This shared task includes three tracks: Strict, Strict-small, and Loose.

The Strict and Strict-small tracks require that submissions are trained exclusively on a fixed dataset, which we provide. The main difference between these tracks is the size of the dataset (∼\sim10M words vs. ∼\sim100M words). Both datasets contain child-directed speech, transcribed speech from multiple sources, children’s books, and Wikipedia, among other datasets. The Strict-small dataset is an approximately 10% uniform subsample of the Strict dataset. See §4 for a full description of the fixed datasets. Winners will be determined based on performance on the shared evaluation set.

The Loose track relaxes these restrictions. Submissions must still be trained on a maximum of 100M words, and will be tested on the shared evaluation set. However, they are permitted to use unlimited non-linguistic data or text which differs from the restricted shared task. Training on additional text is allowed without limits if that text is generated by a model trained following the above restrictions. For this track, winners will be selected holistically based on evaluation performance, relevance to the shared task goals, potential impact, and novelty.

Dataset

We distribute a developmentally plausible pretraining dataset inspired by the input to children.Clicking on the following link will download the dataset (240MB zipped, 700MB unzipped): https://github.com/babylm/babylm.github.io/raw/main/babylm_data.zip Submissions must use only this training data to be considered for the Strict(-small) tracks, but may use different data for the Loose track. The dataset has two key properties:

Under 100M words: Children are exposed to 2M-7M words per year (Gilkerson et al., 2017). Choosing the beginning of adolescence (age 12) as a cutoff, the dataset should be between 24M-84M words.

Mostly transcribed speech: Most of the input to children is spoken. Thus, we include a higher proportion of transcribed speech in our dataset.

The datasets we release are mixed domain, taken from multiple sources. Table 1 summarizes the composition of the datasets.

Evaluation

We will distribute a shared evaluation pipeline based in Google Colab. Colab provides access to relatively small GPUs; this will allow users from various research settings of varying resources to efficiently evaluate their submissions. Our evaluation code will also be public, such that those wishing to use their own computational resources may do so. More details about the evaluation pipeline and the set of tasks will be released subsequently.

The pipeline assumes all models can be loaded and queried in HuggingFace’s transformers library (Wolf et al., 2020).While discouraged, participants whose models are not compatible with the transformers library can still conduct the necessary evaluation through their own pipeline. Additionally, all models must be able to score a sequence—e.g., assign a log-likelihood or pseudo log-likelihood Wang and Cho (2019); Salazar et al. (2020)—and must be able to be fine-tuned to perform classification tasks. Models do not need to be able to generate sequences. Submissions must include model outputs for each of the core evaluations in a format that we specify in our evaluation pipeline.

We choose evaluations that represent the core interests of this shared task, focusing on efficiency and applied NLP, as well as cognitive science, linguistics and language acquisition. Especially good performance in one but not both of these areas may be acknowledged with a special award.

We will also release a series of baseline models with the evaluation pipeline. To train these, we simply take the hyperparameters from a series of established large language models and train them from scratch on our fixed datasets. We use hyperparameters from OPT (decoder-only; Zhang et al., 2022), RoBERTa (encoder-only; Liu et al., 2019), and T5 (encoder-decoder; Raffel et al., 2020). These are not meant to be strong baselines, but rather to provide a naïve starting point for improving language models for this domain.

Submissions

HTML]E9F0E9 What you Need to Submit • A link where we can download the model • A .zip of predictions (from our eval pipeline) • A short description of the approaches taken • If Loose track: a link where we can download any additional data

Although scaled-down pretraining is more accessible to research groups with limited resources, pretraining is still expensive from a computational, energy, and financial perspective. To help groups plan for total costs, we will release an estimate of the resources required to pretrain on 10M words and 100M words. For the Loose track, evaluation of submissions may take into consideration computational efficiency as part of the holistic evaluation.

FAQs

Yes. For example, a single paper can describe models which are submitted separately to the Strict and Strict-small tracks.

Can I submit a paper about my work?

Yes, we encourage all participants to submit their reports, which will be published in the proceedings of CoNLL. You may also describe any additional experiments beyond those required for the shared task evaluation.

Can I submit additional evaluation metrics?

Yes, if you wish to submit your own evaluation metrics, along with model performance, alongside our standardized evaluation results these can be considered as part of the holistic evaluation in the Loose track.

What training regimes are permitted?

For the Strict/Strict-small tracks, any kind of training objective/regime is permitted, as long as the data restrictions are followed. Pretrained models may not be used for any purpose such as reranking or data augmentation.

We do however require for evaluation purposes that the model provides a function to score a sequence—e.g., log-likelihood for autoregressive models or pseudo-log-likelihood for masked language models—without the need for additional fine-tuning.

Are there any limits on hyperparameters?

No. In the Loose track, parameter efficiency and training efficiency may be considered along with other factors in ranking submissions.

Are there any limits on the number of epochs?

No. We put no restrictions on the number of epochs, for several reasons: First, from an engineering perspective, training LMs with SGD tends to require multiple epochs at these scales to achieve peak performance. Second, from a cognitive perspective, humans have a memory of linguistic experience, and can continue to access and learn from these memories. Third, we try not to make a stand on implementations to allow the most freedom for innovation.

Organizing Committee

Questions? Feel free to contact us at the following email addresses: leshem.choshen@mail.huji.ac.il haokunl@cs.unc.edu amueller@jhu.edu alexwarstadt@gmail.com ewilcox@ethz.ch chengxuz@mit.edu

References