TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
Ronen Eldan, Yuanzhi Li
Introduction
Natural language is rich and diverse. It is not only a system of rules and symbols, but also a way of conveying and interpreting meaning . To understand and produce language, one needs not only to master the technical rules of grammar and knowledge of vocabulary, but also to have sufficient factual information and to be able to reason logically and contextually. Therefore, autoregressive language models, which are able to generate coherent English text, must have acquired some degree of these capabilities as well. For example, consider the following incomplete sentence:
Jack was hungry, so he went looking for
To complete this sentence in a sensible way, the language model needs to know that hunger is a state that motivates people to seek food, and that food is a category of things that can satisfy hunger. It also needs to choose a word that fits the syntactic and semantic constraints of the sentence (such as “a snack”), and that is plausible given the situation and the background knowledge.
An example that illustrates the need for reasoning is:
Lily wanted to get either a cat or a dog. Her mother didn’t let her get a dog so instead she
To complete this sentence, the language model needs to invoke reasoning: it needs to apply the principle of disjunction elimination: if Lily wants either a cat or a dog, and she cannot get a dog, then she must choose a cat. It also needs to choose a words that expresses Lily’s intention or action that is coherent with the tone and style of the text.
Language models have been shown to exhibit a range of emergent abilities, such as summarization, arithmetic, translation, and commonsense reasoning, as they are scaled up in size and trained on diverse and large corpora . These abilities suggest that language models are not only learning the surface patterns of language, but also acquiring some degree of semantic and logical understanding of the world and the text. However, it is not clear at what scale these abilities emerge, and how they depend on the model architecture and the data distribution.
Perhaps the most fundamental ability for a language model is to produce coherent and fluent English text, which, as we discussed above, requires not only grammatical and lexical knowledge, but also factual information and contextual reasoning. How well can language models generate text that is consistent, diverse, and meaningful? And what are the minimal requirements for a language model to achieve this ability?
So far, the evidence points to the fact that producing coherent text already requires quite a large scale: small language models (SLMs) are very limited in their performance and capabilities, especially in text generation tasks. For example, models with around 125M parameters such as GPT-Neo (small) or GPT-2 (small) can rarely generate any consistent text beyond a few words even after extensive training on large corpora such as the Pile , Common Crawl or the CC-100 . These models often produce incoherent, repetitive, or nonsensical sentences, and fail to maintain a clear topic or a logical structure across paragraphs . This raises the question of whether the emergence of the ability to speak coherent English requires large models (with hundreds of millions of parameters or more) and complex architectures (with many layers of global attention).
However, it is currently not clear whether the inability of SLMs to produce coherent text is a result of the intrinsic complexity of natural language, or of the excessive breadth and diversity of the corpora used for training. When we train a model on Wikipedia, for example, we are not only teaching it how to speak English, but also how to encode and retrieve an immense amount of facts and concepts from various domains and disciplines. Could it be that SLMs are overwhelmed by the amount and variety of information they have to process and store, and that this hinders their ability to learn the core mechanisms and principles of language?
This raises the question of whether we can design a dataset that preserves the essential elements of natural language, such as grammar, vocabulary, facts, and reasoning, but that is much smaller and more refined in terms of its breadth and diversity. Such a dataset would allow us to isolate and examine the minimal requirements for a language model to generate coherent and fluent text, and to evaluate its performance and capabilities more precisely and fairly. Moreover, such a dataset would facilitate the development and analysis of SLMs, especially for low-resource or specialized domains, where large and diverse corpora are either unavailable or undesirable.
In this paper, we introduce TinyStoriesThe dataset is available on Huggingface named TinyStories., a synthetic dataset of short stories that are intended to contain only words that most 3 to 4-year-old children would typically understand, generated by GPT-3.5 and GPT-4. TinyStories is designed to capture the essence of natural language, while reducing its breadth and diversity. Each story consists of 2-3 paragraphs that follow a simple plot and a consistent theme, while the whole dataset aims to span the vocabulary and the factual knowledge base of a 3-4 year old child
Based on this dataset, our paper makes several main contributions:
Our main contribution is that we show TinyStories can be used to train and evaluate SLMsOur models are available on Huggingface named TinyStories-1M/3M/9M/28M/33M/1Layer/2Layer and TinyStories-Instruct-. We use GPT-Neo architecture with window size 256 and context length 512. We use GPT-Neo tokenizer but only keep the top 10K most common tokens. that are much smaller than the state-of-the-art models (below 10 million parameters with an embedding dimension of 256), or have much simpler architectures (with only one transformer block), yet still produce a diverse set of fluent and consistent stories that are comparable or superior to those generated by larger and more complex models. Moreover, despite of the small size of the models, we still observe an emergence of reasoning capabilities, knowledge of general facts and ability to follow certain instructions.
We introduce a new paradigm for evaluating language models using GPT-4, which overcomes many of the limitations of standard benchmarks.
We show that although the training of generative models on TinyStories can typically be done in less than a day on a single GPU, they still exhibit many behaviors similar to the ones observed in LLMs, such as scaling laws, trade-offs between width and depth, etc. Even with limited computational resources, we are able to conduct extensive experiments to study the effects of different hyperparameters, architectures and training methods on the performance and quality of the models.
We show that the trained SLMs appear to be substantially more interpretable than larger ones. When models have a small number of neurons and/or a small number of layers, we observe that both attention heads and MLP neurons have a meaningful function: Attention heads produce very clear attention patterns, with a clear separation between local and semantic heads, and MLP neurons typically activated on tokens that have a clear common role in the sentence. We visualize and analyze the attention and activation maps of the models, and show how they relate to the generation process and the story content.
To give the reader a first impression of the abilities of models trained on TinyStories, we compare the completion of a 28M parameter model trained on TinyStoriesFor the sake of replicability, most completions which appear in this paper, including this one, were generated with zero temperature. with that of GPT2-XL, which is two orders of magnitude bigger (1.5B parameters), on a sample promptThis prompt was composed manually and then verified to have no 6-gram overlap with the dataset. in Figure 1. We remark that the architectures and training scheme of the models are essentially the same.
Returning to the examples given at the beginning of the introduction, we highlight the completions in Figure 2. Those completions, along with many other examples given throughout the paper, demonstrate that even very small models (2.5M) or models with only one transformer layer are able to attain factual knowledge, and that slightly bigger models are able to express a certain extent of reasoning. The comparison with the performance of GPT2-XL reveals that the TinyStories dataset allows us to exhibit the emergence of these capabilities at significantly smaller scales compared to models trained on large language corpora. More evidence for this will be provided throughout the rest of the paper.
Description of the TinyStories dataset
As mentioned above, the idea behind the TinyStories dataset is to have a corpus that combines all the qualitative elements found in natural language, such as grammar, vocabulary, facts, and reasoning, but is also smaller, less diverse, and more restricted in terms of its content. A natural approach to achieve this is to use the fact that young children qualitatively attain the aforementioned intellectual abilities, albeit with much less exposure to language than an adult human .
To this end, we rely on the latest text generation models by OpenAI (GPT-3.5 and GPT-4) which are able to produce large amounts of synthetic content according to instructions. In particular, we instruct the models to produce content that only uses vocabulary that a typical 3-year-old child would understand. We restrict the content to have the format of short stories in English. The main challenge in using large language models for producing training data is generating a dataset that is sufficiently diverse: prompting those models to produce stories, even if the temperature of generation is set to a high value, will still produce a very repetitive dataset, whose diversity is very far from what is required for training a language model that has a comparable ”understanding” of language to that of children.
In order to address the problem of creating a diverse dataset, we collected a vocabulary consisting of about 1500 basic words, which try to mimic the vocabulary of a typical 3-4 year-old child, separated into nouns, verbs, and adjectives. In each generation, 3 words are chosen randomly (one verb, one noun, and one adjective). The model is instructed to generate a story that somehow combines these random words into the story. As we argue below, this greatly increases the diversity of the dataset, forcing the stories to span the entire vocabulary a child is familiar with, and to include a rich set of ways to combine different concepts. In addition, we constructed a list of possible features a story could have (such as a dialogue, a plot twist, a bad ending or a moral value). For each story we generated a random subset of those features and prompted the model with the extra requirement for the story to have these features.
Here is a sample generation, including the prompt and the story generated by GPT-3.5.
Write a short story (3-5 paragraphs) which only uses very simple words that a 3 year old child would likely understand. The story should use the verb ”decorate”, the noun ”thunder” and the adjective ”ancient”. The story should have the following features: the story should contain at least one dialogue, the story has a bad ending. Remember to only use simple words!
Once upon a time, in an ancient house, there lived a girl named Lily. She loved to decorate her room with pretty things. One day, she found a big box in the attic. She opened it and saw many shiny decorations. Lily was very happy and decided to use them in her room. As Lily was decorating her room, the sky outside became dark. There was a loud thunder sound, and Lily got scared. She ran to her mom and said, ”Mommy, the thunder is so loud!” Her mom hugged her and said, ”Don’t worry, it will pass soon.” But the thunder did not stop. It got louder and louder, and the ancient house started to shake. Suddenly, the roof fell down on the pretty decorations. Lily was sad because her room was not pretty anymore. The end.
Another challenge of natural language generation is to produce texts that are coherent, relevant, and consistent with given instructions or constraints. To evaluate the capability of generative models in terms of instruction-following rather than just content creation, we created a variant of TinyStoriesAvailable on Huggingface under the name TinyStoriesInstruct. (referred to as TinyStories-Instruct) in which each story is preceded by a corresponding set of instructions. The instructions are of four types:
A list of words to be included in the story.
A sentence that should appear somewhere in the story.
A list of features (possible features: dialogue, bad ending, moral value, plot twist, foreshadowing, conflict).
A short summary (1-2 lines) of the story.
Every entry in this dataset has a random subset of the these instructions followed by the story itself. This way, given a model trained on this dataset, we can compose a prompt to reflect any set of instructions and use the model to generate a corresponding story. Here is an example from the dataset:
Summary: Lily and Timmy build a sandcastle together and learn to compromise, but it gets knocked over by a gust of wind. They find beauty in the broken sandcastle and play happily with a butterfly. Features: Dialogue, Foreshadowing, Twist Sentence: One day, she went to the park and saw a beautiful butterfly. Words: disagree, network, beautiful Story: Once upon a time, there was a little girl named Lily. She loved to play with her toys and her friends. One day, she went to the park and saw a beautiful butterfly. She tried to catch it, but it flew away. Lily’s friend, Timmy, came to play with her. They decided to build a sandcastle, but they disagreed on how to build it. Lily wanted to make it tall, while Timmy wanted to make it wide. They argued for a while, but finally agreed to make it both tall and wide. As they were building the sandcastle, they noticed a group of ants working together. They realized that the ants were building a network of tunnels. Suddenly, a gust of wind came and knocked over their sandcastle. They were sad, but then they saw the butterfly again. It landed on the remains of their sandcastle and they realized that it was still beautiful, even in its broken state. They smiled and played together happily.
TinyStories-Instruct was created in the following way: For each story in the original dataset, we already had a list of words and features that were used to create it. In addition, we used GPT-3.5 to create short summaries of our stories, and from each story we extracted a random sentence (which could be any sentence in the story except for the first one). Then we chose a random subset of these and combined them into an instruction, followed by the story itself.
GPT-Eval: A framework for evaluating models trained on TinyStories
While the common paradigms for evaluation of language models usually rely on structured evaluation datasets in the form of a task where the output of the model has to match a given answer, we introduce a new paradigm that is arguably more suitable in this context. Again, we take advantage of existing large language models to evaluate the stories generated by the model. We provide the model with a story’s beginning, taken from a manually-prepared dataset consisting of around 50 prompts, generate a completion using the model, and provide the story’s beginning together with the model’s completion to GPT-4, asking it to grade the completion assignment in terms of grammar, creativity, and its consistency with the beginning of the story. The prompts usually contain a partial sentence that needs to be completed, where we attempted to write them in a way that it would be a grammatically nontrivial task to complete the sentence that was cut short.
As an example, we took the story that appears in the example above and cut it after the word ”loud”. We generated a completion of the above story using a model with roughly 28M parameters (the generation is highlighted in yellow), then we prompted GPT-4 to first provide a verbal assessment and then grade the completion. The transcript of interaction is as follows, with the parts produced by GPT-4 in italic font:
the following exercise, the student is given a beginning of a story. The student needs to complete it into a full story. The exercise tests the studentś language abilities and creativity. The symbol *** marks the separator between the prescribed beginning and the student’s completion:
Once upon a time, in an ancient house, there lived a girl named Lily. She loved to decorate her room with pretty things. One day, she found a big box in the attic. She opened it and saw many shiny decorations. Lily was very happy and decided to use them in her room. As Lily was decorating her room, the sky outside became dark. There was a loud*** thunderclap and a bright flash of lightning. Lily was a little scared, but she knew she had to be brave. She told her mom and dad, and they all went outside to see what was going on. When they got outside, they saw a big storm coming. The wind was blowing hard, and the rain was pouring down. Lily, her mom, and her dad knew they had to find a safe place to hide. They found a big tree and hid under it. The storm passed, and the sun came out again. Lily, her mom, and her dad were all safe and warm inside their ancient house. Please provide your general assessment about the part written by the student (the one after the *** symbol). Is it gramatically correct? Is it consistent with the beginning of the story? Pay special attention to whether the student manages to complete the sentence which is split in the middle by the separator ***.
The student’s completion of the story is mostly consistent with the beginning of the story. It maintains the focus on Lily and her family, and the sentence split by the separator is completed correctly. However, the student’s addition does not fully integrate the shiny decorations found in the attic, which were a significant part of the beginning.
The grammar is generally correct, but there are a few minor errors: .
Overall, the student’s completion of the story demonstrates adequate language abilities and creativity, but could benefit from better integration of the shiny decorations and minor grammar improvements.
Now, grade the student’s completion in terms of grammar, creativity, consistency with the story’s beginning and whether the plot makes sense. Moreover, please provide your best guess of what the age of the student might be, as reflected from the completion. Choose from possible age groups: A: 3 or under. B: 4-5. C: 6-7. D: 8-9. E: 10-12. F: 13-16.
Grammar: 8/10, Creativity: 7/10, Consistency: 7/10, Age group: E (10-12)
To perform the full evaluation, for each of the manually-constructed prompts in the evaluation set, we use the trained model to generate 10 completions with temperature 1. We average the GPT-4 evaluation scores of all the completions. Figure 3 provides an example of the evolution of the different scores as the training progresses, along with the training and evaluation loss, for several model sizes. Figure 4 shows how the different scores vary with different model sizes and architectures, after a fixed number of training steps.
Our evaluation method for models trained on TinyStories-Instruct also relies on GPT-4. We have implemented a corresponding evaluation pipeline, with an evaluation dataset containing 50 different sets of instructions to create a story (we verified that these are disjoint from the training set). In the scoring phase, we provide GPT-4 with both the instructions and with the generated story. We prompt GPT-4 to base the consistency score on the extent to which the story accurately reflects the given instruction. In addition, we added a Plot category that reflects the extent to which the plot is coherent. Figure 5 illustrates the whole pipeline that combines the generation of the story by our model, and its evaluation by GPT-4. Scores assigned to models of different sizes appear in the two right-hand columns of the table in Figure 4.
Our proposed evaluation method gives a way to obtain a more fine-grained assessment of the model, due to which we can draw conclusions regarding the dependence of different types of capabilities on the size and architecture of the model. While all the evaluation scores are consistently increasing with the decrease of evaluation loss, a more careful scrutiny of the results reveals the following:
Figure 3 suggests that shallower models perform better in terms of grammar compared to content consistency, meaning that model depth is more important for keeping consistent with the content than for generating syntactically correct language (we provide additional evidence for this in the next section).
In the same figure, we observe that the score for grammar plateaus at an earlier stage than the other two scores. Furthermore, in Table 4, we also see that while grammar can be mastered by relatively small models, consistency and creativity only emerge at a larger size.
Table 4 further suggests that the ability to generate a completion that is consistent with the beginning of the story emerges when the hidden size of the model increases from 64 to 128.
We also see that the largest model that we have trained on TinyStories (with roughly 80M parameters) reaches almost perfect scores in terms of grammar and consistency. However, it falls short of GPT-4’s abilities in terms of creativity quite significantly, suggesting that creativity continues to improve more substantially with the sizes of the model and dataset, compared to grammar and consistency.
The right-hand columns of Table 4 suggests that the models that have only 1 layer seem to struggle quite substantially with following instructions (which likely heavily relies on global attention), and 2 layers seem to be sufficient for a certain extent of instruction-following. Comparing the ”Instruct” and ”Plot” scores we also see that the quality of instruction-following depends more heavily on the number of layers, in comparison with the coherence of the plot for which the hidden dimension is more important.
The performance of small models trained on TinyStories
In this section, we give some initial examples that illustrate how TinyStories gives rise to models of very small size that can generate coherent language and exhibit common-sense knowledge as well as certain extent of reasoning capabilities. We also provide evidence that the generated content is truly diverse, refuting the possibility that the models simply output content that has been ”memorized”.
Throughout the section, we work with several architectures of models whose size ranges between roughly 1M and 35M parameters, and whose number of layers range between 1 and 8 layers. All of the models can be trained on a single V100 GPU within at most 30 hours.
To get a glimpse into the capabilities of (very) small models trained on TinyStories, let us first inspect a few examples. We start with the following story, taken from the training set.Below, there will be many other examples of completions of stories from outside of the training set.
Sara and Ben are playing in the snow. They make a big snowman with a hat and a scarf. They are happy and laugh.
But then a big dog comes. The dog is angry and barks. He runs to the snowman and bites his hat. Sara and Ben are scared and cry. ”Go away, dog! Leave our snowman alone!” Sara shouts. But the dog does not listen. He bites the scarf and the snowman’s nose. He shakes his head and makes the snowman fall.
Sara and [ Ben run to their house. They slam the door and lock it. They are safe and hug each other. ”Mom, mom, a big dog broke our snowman!” Ben says. Mom comes and hugs them. She is sorry and kisses them. ”Don’t worry, my loves. We can make another snowman tomorrow. But now you need to get warm and have some hot chocolate. You are cold and wet.” She takes them to the kitchen and makes them hot chocolate. She gives them cookies and marshmallows. She tells them to drink and eat in order. Sara drinks first, then Ben. Ben eats first, then Sara. They feel better and smile. They thank mom and tell her they love her. They escape from the big dog and the cold. They are happy and warm. ]
We now cut the story roughly in the middle, removing the part inside the square brackets and generate an alternative ending with several different models.
To set a baseline, we begin with the completion generated by GPT2-XL, whose size is roughly 1.5 Billion parameters:
…Sara and Ben cry and look at the snowman in the snow. Ben asks to do something. Sara tells him to go away. Her hand tightens on the little red button. The snowman is covered in fur. It is red. It is real hair. It is a real snowman. It looks just like the ones they used to make. Ben has a big smile on his face. He looks at the face he has made. And he walks over and kisses the eyes. He is happy to be alive. He has done what he was going to do. He had saved the world.
Now, let us provide the generation given by a model trained on TinyStories, whose size is about 2.5 million parametersThis example was generated with temperature 0.5.:
…Sara and Ben are very sad and angry. They cry and shout at the dog. ”Bad dog! Bad dog! Go away!” Ben says. ”Go away, bad dog! Leave us alone!” Sara says. The dog does not go away. He wags his tail and licks their faces. Sara and Ben feel sorry for the dog. They want to make him happy. ”Maybe we can make a new snowman for him,” Sara says. ”Maybe,” Ben says. They hug the dog and give him a name. The dog wags his tail and licks their faces. ”Maybe he likes us,” Sara says. ”Maybe he likes us,” Ben says. The dog wags his tail and licks their faces. He is happy and friendly.
First of all, we remark that the above generation is very different from the original ending of the story, which provides a first evidence that it hasn’t simply been memorized (we discuss this much more thoroughly later on). We also see that the generated content is consistent with the beginning in terms of plot, characters and events. Arguably, this continuation is of higher quality than the one provided by GPT2-XL despite the size of the model being almost 1000 times smaller.
However this raises the question whether the small models are effectively memorizing a small number of templates and the generation is very limited in terms of diversity, which would be a valid concern at this point. We systematically address this later on, in Section 4.4. By testing the overlap of the generated stories with the training set and by considering out of distribution generation, we are led to the conclusion that those models do have substantial diversity, and do not rely on memorization.
Next, we consider the completion by a model whose size is roughly 21M parameters, but only has one transformer layer (and consequently only one layer of global attention):
…Sara and Ben are very sad. They cry and hug each other. They wish the dog would go away. Then they hear a voice. It is their mom. She comes out of the house and sees them. ”What happened here?” she asks. ”Why are you crying?” Sara and Ben tell her what happened. They tell her about the dog and the snowman. Mom hugs them and says, ”Don’t worry, my loves. The dog didn’t mean to hurt you. He was just curious and playful. He didn’t mean to hurt you. He just wanted to play.” Sara and Ben feel better. They say sorry to the dog and the snowman. ”Can we still play with the dog?” Sara asks. Mom smiles and nods. She says, ”Yes, you can. But you have to be gentle and careful. And you have to share the snowman with us.” Sara and Ben nod and say, ”OK, Mom. We will.” They go inside and have some hot chocolate. They are warm and yummy. They are happy.
While this completion is far from perfect, it is (arguably) consistent with the beginning in the sense that the model successfully captured the tone and the main elements of the plot.
We acknowledge that the two examples above are not the worst-case scenarios for the models, but they are not extremely rare either. Models of this size can sometimes produce less coherent or plausible completions, but they can also often match or surpass the quality of the ones shown here. However, if we increase the number of parameters by an order of magnitude, we observe that the models consistently generate coherent and relevant continuations. For the sake of replicability, examples from this point on will generated at zero temperature, and in addition, the models parameters are provided as supplemental material. For the sake of replicability, examples below are all generated at zero temperature, and in addition the models parameters are provided as supplemental material.
In order to give the reader an impression of the dependence of the quality of completions on the size of the model, Figures 6, 7 and 8 each provide different completions for one prompt given by models of different sizes and depths. Each table represents a different prompt, which we have manually composedWe manually verified that the dataset does not contain any entries which are similar or close to these prompts..
We see that the quality of generation clearly improves as a factor of size, and appears to be consistent with the grades given by the GPT-4 evaluation. The smaller model (64_8) can barely produce a completion which looks coherent with the beginning of the story, and often repeats itself or makes no sense. As the size increases, the models become more and more coherent, and the grammar becomes better. The models can also generate more diverse and creative endings, and use more details and emotions.
We can also notice that models with a small number of layers have a hard time staying in context, even if they do manage to produce syntactically correct English. This suggests that the model lacks the ability to capture the long-term dependencies and the structure of the story. On the other hand, models with more layers can better maintain the consistency and the logic of the story.
An interesting observation is that in Figure 7, even though the completions are generated by different models, they begin in a very similar way (all completions have to do with a little girl coming by and talking to the pumpkin). We point out that the reason for this seems to be that the completions are generated with temperature 0. Roughly speaking, this gives rise to the ”most likely” completion. In order to demonstrate that the model is capable of generating a more diverse set of endings to the story, we added a completion with a non-zero temperature. It appears, however, that the quality of completion slightly decays when increasing the temperatureWe do not present more evidence for this claim as it goes beyond the scope of the paper..
2 Knowledge, reasoning and context-tracking
Next, we assess the capabilities of the different models on three additional types of prompts:
Factual prompts, which test the models’ knowledge of common sense facts.
Reasoning prompts, which test basic reasoning abilities, such as cause and effect and elimination.
Consistency (context-tracking) prompts test the models’ ability to maintain coherence and continuity with the given context, such as the names and actions of the characters, the setting and the plot.
We report the generated continuations for each model and prompt in three tables (Figure 9, Figure 10 and Figure 11), and color-code them according to their success (green), failure (red), or partial success (yellow).
The results show that as the embedding dimension and the number of layers increase, the performance in regards to all three categories improve. The models with higher embedding dimensions and more layers tend to generate more accurate, relevant, and natural continuations, while the models with lower embedding dimensions and fewer layers tend to generate more nonsensical, contradictory, or irrelevant continuations. For example, the model with 1M parameters and 8 layers fails to answer any factual prompt correctly, and often generates sentences that do not make sense or do not follow the grammar. The model with 33M parameters and 8 layers, on the other hand, answers most prompts, from all three categories, correctly. Comparing to the completions given by GPT2-XL (right hand column), we see that despite its much larger size, its performance in all three categories is worse than some of our models.
One interesting finding is that knowledge of facts seems to rely more on the embedding dimension, whereas for context-tracking the number of layers is more important. For example, the model that has only 1 layer does not get any consistency prompt right, but does get some facts right, whereas the model with embedding dimension 64 does not get any fact right, but manages to maintain consistency several times. This suggests that the embedding dimension is more crucial for capturing the meaning and the relations of words, while the number of layers is more crucial for capturing long-range dependencies in the generation.
3 Instruction-following examples and out-of-distribution generation
Table 12 provides an example of the generation of different models trained on the TinyStories-Instruct dataset, together with the evaluation scores given by GPT-4. As the model size increases, we see an improvement both its ability to follow instructions and to generate a coherent plot.
This dataset also enables us to test whether our models have a reasonable out of distribution performance. Recall that in each entry of TinyStories-Instruct, the instructions are created as a (random) combination of possible types of instructions (words to use, summary, prescribed sentence, features). We created another variant of the TinyStories-Instruct (called TinyStories-Instruct-OOD) where we disallowed one specific combination of instruction-types: The dataset does not contain any entry where the instruction combines both the summary of the story and the words that the story needs to use (we chose this particular combination because in a sense, it is the most restrictive one). We then tested whether models trained on this variant would be able to produce stories that follow these two types of instructions combined. An example is provided in Figure 13, for a model with 33M parameters. We see that, perhaps somewhat surprisingly, the model is able to follow these two types of instructions simultaneously even if it has never been trained on such a task.
4 Diversity of the content generated by the model
One of the main challenges of text generation is to produce diverse and creative texts that are not just repetitions or variations of existing texts. Our small models can generate coherent and fluent English text, but this would not be very impressive if they were simply copying or paraphrasing large portions of the dataset. Therefore, in this section, we aim to address this concern. We will provide several methods and metrics that show that the models can generate diverse texts that are not similar to any story in the dataset, and that they can adapt to different instructions and contexts.
To evaluate the diversity of the content generated by the models, we first need to define what we mean by memorization, and what kinds of memorization we want to avoid or detect. We classify three levels of memorization as follows:
Exact memorization: This is the simplest and most obvious form of memorization, where the model simply copies an entire story or a large portion of it from the dataset, without changing anything. This can be easily detected by checking the similarity or the hash of the generated story with the stories in the dataset.
Simple template matching: This is a slightly more sophisticated form of memorization, where the model changes some names or entities in a story from the dataset, but keeps the rest of the story the same. For example, the model might change the names of characters, or the location of the story, but keep the plot and the events the same. This can be detected and prevented by measuring the overlap of words and n-grams between the generated story and the stories in the dataset.
Complex template matching: This is the most subtle and difficult form of memorization, where the model follows a more abstract pattern or structure from the dataset, keeping the general plot but changing the details and the specifics of the story. This is almost impossible to quantify, as it requires a deeper understanding and analysis of the content and the meaning of the stories, and how they relate to each other.
We claim that our models are not doing exact memorization or simple template matching, as evidenced by the methods and metrics we use to evaluate the diversity of the content generated by the models. We rely on several approaches:
Manual inspection: We generate completions for a range of human-constructed stories. We inspect the stories generated by the models and check that they are not copies or close modifications of the stories in the dataset.
Completion of training stories: We take stories from the training set, truncate them in the middle and generate alternative completions with our models. We then compare the completions with the original stories. We observe that the completions are typically very different from the original stories, and often introduce new characters, events, or twists. This is shown in Figure 14.
Diversity of instructions: Recall that in the TinyStories-Instruct dataset, we provide a set of instructions in the form of summaries or words contained in the stories, followed by the stories themselves. We can then change the instructions, verify that the combinations do not appear in the dataset and see how the models adapt to the new instructions. We find that the models can generate diverse stories that follow the instructions, even if they are novel or challenging, such as requiring the model to fit unlikely words into the story or adding features such as a plot twist or a bad ending.
We measure the diversity of the stories quantitatively using word and n-gram overlap. We inspect the overlap of words and n-grams between different stories generated by the models, and compare them with the overlap in the dataset. We find that the models’ generations have a very low overlap with the dataset, indicating that they are not repeating the same words or phrases. We use the standard Rouge score, for the source text with -gram respectively, the rouge precision score is defined as:
The Rouge precision score measures how many -grams in is included in that of . The final Rouge score (fmeasure) is given as:
We perform the following experiment: We randomly pick stories from the training dataset, we cut each story in the middle, keeping roughly the first , and use it as a prompt. We ask the model to generate a completion from each prompt. Let be the generated completions and be the original completion, we measure:
How much of the new generation is contained in the original story (Figure 14), meaning:
How similar are the generated stories to each other (Figure 15), meaning:
To what extent are the k-grams in the generated story copied from the training dataset (Figure 16). More precisely, we take as the entire training corpus, for each we measure
In other words, for each -gram generated by the model, we measure the frequency that it appears in the original training dataset, where means that the -gram never appears in the training dataset.
How similar is the generated story to the closest point, in terms of Rouge precision score, in the entire dataset. Let be all the stories in the training dataset, in Figure 17, we compute
For the sake of getting a more concrete impression about how different the model completions are from the original ending of the story and from other stories in the dataset, in Figure 18 we provide one example of the original story, the alternative completion by our model together with its closest point in the training dataset.
The above points towards several findings:
When the model generates stories using a diverse set of prompts, it ends up with a diverse set of completions.
When completing stories from the dataset, the completions usually turn out to be very different than the original story.
Typical -grams in generated completions rarely appear in the dataset, for values of as small as or .
The closest point in the dataset to each generated completion is typically still quite far from it.
All the above, taken together with the ability of models trained on TinyStories-Instruct to successfully follow sets instructions which we can easily be verified to be disjoint from the dataset (for example, combinations of words can be checked), provides strong evidence that our models produce genuinely novel and diverse stories, rather than simple variations of existing stories.
We remark that nevertheless, we are not able to completely rule out the possibility that the models perform complex template matching, as it is hard to define and measure what constitutes a novel plot or a novel story. We acknowledge that this is a limitation of our evaluation. Another possibility is that the stories in the dataset essentially span the entirety of support of the distribution in the (weak) metric of complex template matching.
Interpretability
Understanding the inner workings of deep neural networks and language models in particular is a major challenge in this field of study. For example, it is often difficult to assign a specific function to a given component of a neural network. This may be because, contrary to our intuition based on human-designed programs, the network components may not have distinct roles, but rather interact in a complex and messy way. In this section, we present some preliminary evidence that training smaller models on TinyStories leads to higher interpretability, suggesting that when networks are constrained in size, we may be able to gain some insights into their internal mechanisms. We focus on two aspects of the model: the attention heads and the neurons in the MLP.
As this is not the main focus on our paper, this section is by no means exhaustive and much more work is required in order to reach more conclusive findings. Rather, we only give some preliminary evidence which may hopefully motivate future work.
In the study of attention heads, we take advantage of the fact that we were able to train a very shallow model (having only one transformer block) which still manages to generate meaningful text. Since the model has only one layer, the attention heads are directly responsible for generating the output tokens, and thus they may have more interpretable functions than in deeper models. We use the method of Voita et al to analyze the attention patterns of the heads and classify them into different types, such as positional, syntactic, or semantic. We also use the method of Clark et al to visualize the attention maps of the heads and inspect their behavior on specific examples.
Our findings suggest that the attention heads exhibit diverse and meaningful functions, such as attending to the previous word, the subject of the sentence, the end of the sentence, or the main topic of the story. We also observe that some attention heads specialize in generating certain types of words, such as nouns, verbs, or punctuation. These results suggest that the attention heads learn to perform different linguistic tasks and capture different aspects of the stories.
We also give some initial evidence that in smaller models, some neurons in the MLP have roles that are interpretable by humans. We use the method similar to to identify the most influential tokens in the MLP for each neuron. We find that some neurons are activated on words that have a specific role in the sentence (such as the subject or the action), or in the story (such as the introduction of the protagonist). These findings suggest that the neurons in the MLP learn to encode different semantic and stylistic information and influence the generation process.
1 Interpreting the role of different attention heads
To understand the model’s attention pattern after training, we use a 1-layer model with hidden dimension 1024 and 16 attention heads that was trained on TinyStories. We visualize the attention patterns that it produces when processing the following paragraph (the bold form is the prompt, the highlighted text is generated by the model):
One day, Lucy asks Tom: ”I am looking for a banana but I can’t find it”. Tom says: ”Don’t worry, I will help you”. Lucy and Tom go to the park. They look for the banana together. After a while, they found the banana. Lucy is happy. She says: ”Thank you, Tom. You are a good friend.” Tom: ”You are welcome, Lucy. I am happy to help you. Let’s eat the banana together!”
There seems to be a clear separation between heads with attention pattern based mainly on the distance between tokens, and heads whose attention pattern has a stronger dependence on the semantic meaning:
Out of the 16 attention heads, we observe multiple positional-based attention heads, such that each token attends to tokens with a prescribed relative distance. Different heads are associated with different distances.
We also observe that there is (1). one head that the word “the” and “a” all attend to the word “banana”, interestingly, the “the” at “the park” also attends to “banana”, but the model still manage to generate “park”, which is the consistent completion. (2). Another attention head gives a pattern where the tokens “the” and “a” all attend to “park”. (3). There is third head that most of the words attend to the name of “Tom” and “Lucy”.
We remark that it makes sense that the generation of words like “the”, “a”, “and” or “,” would be induced by distance-based, local attention heads, since those are tokens with a grammatical role which depends on the short-range interactions within a single sentence. On the other hand, the main entities in the story such as “banana”, “park”, “Lucy” and “Tom” cannot usually be predicted (as a next token) only based on the neighboring tokens, which is why the model needs to use semantic attention heads for their generation.
2 Interpreting the roles of different Neurons
In order to examine whether neurons have meaningful roles, we follow , and visualize the most significant tokens for each neuron. More precisely, we take a collection of 20 stories (about 8,000 tokens) from our dataset. We take a model that was trained on TinyStories, we pick a transformer layer, and from the MLP associated with it we pick one coordinate in its intermediate layer. We refer to such a choice as a neuron. We process the collection of stories with the model to obtain their internal representations, which gives us an activation value for each combination of token and neuron. Then, for each neuron we look at the tokens with highest activations from the entire collection. We highlight those tokens in red (and present them along with the sentence they are contained in). We repeated this for two models: a small model of hidden dimension 64 and 1M parameters, trained on TinyStories (Figure 21), and on GPT2-XL (Figure 22).
In the 1M-parameter model trained on TinyStories, Figure 21 first presents the activated tokens for the first two neurons in the before-last layerThe rationale behind choosing the penultimate layer is that tokens have already been processed by most layers at this point. We take the before-last rather than the last layer since the hidden representation in the last layer only has only the role of predicting the next token, so information may be lost at that point.. Note that, since the architecture is invariant to permutations between neurons, taking the two first neurons is the same as taking an arbitrary choice of two neurons, the point being that these neurons are neither unique nor have been cherry-picked. We see (top row of the figure) that each of those neurons is activated on tokens with a common role (one is activated on pronouns which are also the subject in the sentence, and the other is activated on the action in the sentence). In addition, we present the activated tokens for the first neuron in another layer (layer 6), where the neuron is activates only on adjectives. Finally, we picked the neuron which has the largest activation values over all combinations of token and neuron. This neuron (depicted in the bottom right) seems to have the role of identifying the first time that the protagonist of the story is presented.
For comparison, Figure 22 presents the activated tokens for first two neurons of layer 12 for GPT-XL, a much larger neural network. In this case, none of the two neurons seem to have an apparent role.
Exploring architectures and hyperparameters for NLP with TinyStories
One of the main challenges in developing large language models (LLMs) comes from the high computational cost involved in training. Finding the best architectures, training algorithms and hyperparameters for LLMs requires a lot of resources and experimentation. Therefore, it would be useful to have a smaller and simpler dataset that can still capture some of the basic capabilities of LLMs, and allow us to study how different design choices affect their performance. TinyStories is such a dataset, as it enables us to train and evaluate LMs that are orders of magnitude smaller than the state-of-the-art models, yet still have the basic capability of producing coherent text.
In this work, we take the first steps towards using TinyStories as a testbed for exploring architectures and hyperparameters for NLP. We show that our small models exhibit some similar patterns to the ones observed in LLMs in certain aspects. In particular, we investigate two questions: how to balance model size and learning budget for a fixed amount of training flops, and how to choose the number of attention heads for a given model width and depth.
For a fixed amount of training flops, there is a trade-off between the size of the model and the number of training steps (the total number of flops is the product of both). Previous works have shown that there is a polynomial scaling law between model size and learning budget for LLMs, i.e., the optimal model size for a given amount of flops is proportional to the flops raised to some power . However, these works used different ranges of model sizes (from a few million to tens of billions of parameters) and found different values of (around 0.7 and 0.5, respectively). A natural question is whether this scaling law is universal or depends on the dataset. Our dataset allows us to conduct a similar experiment but with much smaller models and flops. Surprisingly, we find evidence for a polynomial scaling law as well, which suggests that there might be a universal phenomenon here.
We train models of various sizes and architectures on TinyStories. For each amount of flops, we select the model and the number of training steps that achieve the lowest validation loss among the possible combinations. We vary the number of layers from and the hidden dimension from . The result is shown in Figure 23. Although the number of points may be a bit small for the data to be very conclusive, the plot points to a polynomial dependence.
Another design choice for transformers is the number of attention heads for each layer. It is not obvious how the number of heads affects the performance of the model, given a fixed model width and depth. Our results, shown in Figure 24, suggest that in the regime where the number of heads is small, increasing it improves the performance of the model across all metrics.
Related Works
Generative language models (LMs) have achieved impressive results in various natural language processing tasks, such as text summarization, dialogue generation, and story completion. However, most of these models are very large, with hundreds of millions or even billions of parameters, which poses significant challenges for training, inference, and deployment. For example, GPT-3 , one of largest LM to date, has 175 billion parameters and requires hundreds of petaflops of compute to train. Smaller models, such as GPT-2 small with 125 million parameters, can hardly generate coherent and consistent sentences beyond a few words, even after extensive pre-training on large corpora .
Several methods have been proposed to compress or distill large LMs into smaller ones, such as knowledge distillation , pruning , and quantization . However, these methods are much more effective for BERT-like models , which are designed for masked language modeling and downstream classification tasks, than for GPT-like models, which are designed for autoregressive language generation .
Another challenge for generative LMs is the evaluation of their outputs. Unlike BERT-like models, which can be fine-tuned and evaluated on downstream tasks with labeled data, GPT-like models are more difficult to measure in terms of how well they can ”speak and understand natural language”. Most existing benchmarks for generative LMs, such as LAMBADA , CLOZE , TriviaQA , and Winograd Schema Challenge , require the models to produce a single word or a short phrase as the answer, which does not capture the richness and diversity of generating natural language. Moreover, these benchmarks are often limited by the size and quality of the datasets, the ambiguity and subjectivity of the answers, and the lack of human evaluation. Larger and more diversed datasets such as the BigBench are simply way too complicated for SLMs. Some other benchmarks, such as WikiSQL , have a more structured output format, which makes them easier to evaluate, but also less representative of natural language generation.
Our work is also beneficial to the theoretical analysis of transformer models and their learning process. Most of the existing theory works focus on models with one transformer block, which are easier to analyze than models with multiple blocks. For example, Voita et al showed that one transformer block can learn to perform different linguistic tasks depending on the position of the self-attention layer. Li et al shows a transformer block can encode topical models. Jelassi et al shows one transformer block can encode patch associations. Our work provides empirical evidence that one transformer block can also generate diverse and consistent stories, which suggests that the transformer architecture has a strong expressive power even with a small number of parameters and layers.
Conclusion
In this work, we have presented TinyStories, a synthetic dataset of short stories that only contain words that a typical 3 to 4-year-olds usually understand, generated by GPT-3.5 and GPT-4. We have shown that TinyStories can be used to train and evaluate small language models (SLMs) that are much smaller than the state-of-the-art models, yet still produce fluent and consistent stories with several paragraphs that are diverse and have almost perfect grammar, and demonstrate reasoning capabilities.
While large models trained on the huge and diverse language corpuses on the internet exhibit very impressive capabilities, those datasets appear to be too large for SLMs to capture the complex aspects of language. In this work we have argued that TinyStories enables us to observe and study the emergence of capabilities such as generation of coherent text, reasoning and instruction following in LMs on a much smaller scale, in terms of the size of both model and dataset. By training SLMs on our dataset, we have also observed many behaviors similar to LLMs such as scaling laws, trade-offs between width and depth, etc. Moreover, we have shown that the trained SLMs have much higher interpretability than larger ones, and that we can visualize and analyze their attention and activation patterns to understand how they generate and comprehend stories.
We provided evidence to the fact that the models trained on TinyStories are able to produce genuinely new stories, rather than just copying chunks of text the dataset. It remains a challenge, however, to assess the true extent of the ”creativity” of our models, and to which the models reflect a certain ”understanding” (on a very low level of course) of the stories that they produce as opposed to just template matching to create a plausible continuation. We hope that this dataset can be used in future works to obtain insights about the degree of creativity of language models.
We have also introduced a new paradigm for the evaluation of language models, which uses GPT-4 to grade the content generated by these models as if those were stories written by students and graded by a (human) teacher. This new paradigm overcomes the flaws of standard benchmarks, which often require the model’s output to be very structured, and moreover provides a multidimensional score for the model, providing scores for different capabilities. We believe that this paradigm can be useful much beyond TinyStories.
Finally, we have presented initial findings which point to the roles of width vs. depth in the intellectual capabilities of generative networks, which suggest that width is more important for capturing factual knowledge whereas depth is more important for contextual tracking. Moreover our findings suggest that in terms of emergence, grammatic and syntactic abilities appear earlier than the ability to produce consistent text, which in turn appears ahead of ability to generate content that would be considered as creative. These preliminary findings are only suggestive (and have not been the main focus of this work) but they show how our dataset and evaluation paradigm can enable more fine-grained analysis of the emergence and evaluation of various language capabilities in generative models.
We hope that TinyStories can facilitate the development, analysis and research of LMs, especially for low-resource or specialized domains, and shed light on the emergence of language capabilities in LMs. A general question that arises from this work is whether synthesizing a refined dataset can be beneficial in training networks for practical uses. For example, perhaps it is possible to train a customer service chatbot by synthesizing a large dataset of hypothetical calls.