Art or Artifice? Large Language Models and the False Promise of Creativity

Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu

Introduction

Recent work has explored the potential of using large language models (LLMs) as assistants in creative writing tasks, from stories to screenplays (Mirowski et al., 2023; Yang et al., 2022; Lee et al., 2022; Ippolito et al., 2022; Chen et al., 2023b). An important aspect of this research is to investigate LLMs’ capabilities in terms of their ability to both generate creative content and assess whether a piece of writing is creative, with the ultimate goal of informing interaction design. However, evaluating objectively the creativity of a piece of writing is challenging.

We propose a protocol to evaluate creativity as product grounded in a widely accepted protocol that evaluates creativity as process — the Torrance Tests of Creative Thinking (TTCT) (Torrance, 1966). Based on Guilford’s work on divergent thinking (Guilford, 1967), TTCT measures creativity as a process by testing participants’ abilities in dealing with unusual uses of objects, specific situations, or impossibilities. TTCT is centered around evaluating four dimensions of creativity: fluency (the sheer volume of meaningful ideas produced in reaction to a given stimulus), flexibility (the diversity of categories within the responses), originality (the uniqueness or novelty of answers) and elaboration (the depth or granularity of details within the responses). While the direct application of TTCT might not be possible across diverse creative domains (Amabile, 1982; Baer and McKool, 2009), its four fundamental dimensions have proven adaptable (Trisnayanti et al., 2019; McIntyre et al., 2003; Bourdeau et al., 2020). Thus, based on TTCT and using the Consensual Assessment Technique (CAT) (Amabile, 1982) we design the Torrance Tests for Creative Writing (TTCW) to evaluate creativity as product (Section 3 Design Principle 1 and 2; Figure 1). The Consensual Assessment Technique states that the most valid assessment of the creativity of an idea or creation in any field is the collective judgment of experts in that field. Thus, to design TTCW, we asked 8 creative writing experts in a formative study (Section 4) to propose creativity measures for the evaluation of short fictional stories aligned with the four Torrance dimensions. This resulted in 14 binary tests organized across the four original Torrance dimensions of fluency, flexibility, originality and elaboration (Section 4.2).

To empirically validate TTCW as an evaluation protocol for creativity as product (fictional short stories), we build a benchmark consisting of 48 short stories: 12 stories written by professionals, and 36 by three top performing LLMs (ChatGPT (OpenAI, 2022), GPT4 (OpenAI, 2023) and Claude 1.3 (Anthropic, 2022)) with 1400 words on average per story (Section 5; Figure 1). We recruit a new set of 10 experts, again adhering to CAT, to administer the 14 tests of TTCW on each story, collecting 3 evaluations per story. In this TTCW implementation with experts as assessors we aim to answer three research questions:

Is the TTCW-based creative evaluation consistent and reproducible? In other words, is there agreement among expert annotators when they perform tests on the same stories? Based on a total of 2,000+ expert-administered tests, our findings reveal that experts reach moderate agreement on average (Fleiss Kappa 0.41) across the 14 TTCW tests, and reach strong agreement when considering the aggregated tests (Pearson correlation 0.690.69). This confirms the validity of TTCW for the evaluation of creativity in fictional short stories.

Are the human-written stories more likely to pass individual TTCW than LLM-generated stories? If so, which tests demonstrate the most significant gaps? The analysis of the conducted tests reveals that the 12 expert-written stories pass an average of 84.7% of the tests, confirming that although an expert story must not pass all tests to be deemed creative, experienced writers typically produce artifacts that pass the majority of the tests. In comparison, LLM-generated stories pass many fewer tests on average, from 9% for ChatGPT-generated stories, up to 30% for Claude-generated stories. In other words, LLM-generated stories are three to ten times less likely to pass individual TTCW tests compared to expert-written stories, revealing a wide gap in the evaluated creativity of LLM-generated content.

Which LLMs perform better in TTCW evaluation, and are there specializations observed, with some different LLMs performing better on different Torrance dimensions? Besides the general gap, the granularity of the TTCW reveals that individual LLMs differ in abilities, with GPT4 more likely to pass tests associated with Originality, and Claude V1.3 more likely to pass tests in Fluency, Flexibility and Elaboration.

Prior work has argued that even though LLMs might not be able to directly generate a creative piece of writing, they might be utilized in providing feedback to authors during their writing process (Ippolito et al., 2022). Thus, we perform a study of TTCW implementation with LLMs as assessors (Section 6; and Figure 1) to understand whether LLMs can be used to assess creative writing. We expanded each test into a detailed prompt, and measured whether LLM assessments correlate with collected expert judgments. Our analysis reveals that for the most part, LLMs are not capable of administering the TTCW tests, as the three LLMs we experiment with achieve correlations with experts that are close to zero.

In summary, our work makes the following contributions:

We adapt the Torrance Test for Creative Thinking (TTCT), a protocol for evaluating creativity as a process, and align it for the evaluation of creativity as a product particularly focusing on short stories. Using the Consensual Assessment Technique, we design 14 tests called the Torrance Test for Creative Writing (TTCW) based on the four original Torrance dimensions of fluency, flexibility, originality, and elaboration,

We experimentally validate the TTCW through an assessment of 48 stories involving 10 participants with expertise in creative writing, finding that they reach moderate agreement when administering individual tests, and strong agreement when evaluating all tests in aggregate.

We study the abilities of LLMs to generate stories that pass/fail the TTCW tests and their ability to reliably assess the creativity of stories following the TTCW framework through correlation with human judgments. Our findings show that LLM-generated stories are three to ten times less likely to pass TTCW tests compared to expert-written stories, as well as the fact that current state-of-the-art LLMs are not yet capable of reproducing expert assessments when administering TTCW tests. To enable future research in this fast-evolving domain, we release the large-scale annotation of 2,000+ TTCW assessmentshttps://github.com/salesforce/creativity_eval, each accompanied with a natural language expert explanation.

Finally, we discuss how creative writing experts can distinguish between AI vs. human written stories and how future work can use our evaluation framework for building rich interactive writing support tools.

Related Work

In prior work, Baer (2014) argued how divergent thinking remains the most frequently used indicator of creativity in both creativity research and educational practice, and divergent thinking theory has a strong hold on everyday conceptions of what it means to be creative. Along the same lines, Kaufman et al. (2008) further discusses and evaluates common creativity measures such as divergent thinking tests, consensual technique, peer/teacher assessment, and self-assessment. Silvia et al. (2008) examined the reliability and validity of different subjective scoring methods for divergent thinking tests and introduced a new Top 2 scoring method that involves participants selecting their most creative responses, and demonstrates that this method yields reliable scores with a small number of raters. Plucker et al. (2010) discuss key issues and methods in creativity assessment including reliability, validity, bias, and use of assessments. Beaty and Johnson (2021) explored the use of automated scoring via semantic distance, using natural language processing in assessing the quality of ideas in creativity research, demonstrating its strong predictive abilities for human creativity and novelty ratings across various tasks, thus addressing the labor cost and subjectivity issues in traditional human-rating methods.Like many of the prior works our research centers around grounding creativity evaluation through divergent thinking. In particular, we rely on the Torrance Tests of Creative Thinking (TTCT) as a foundation for measuring creativity.

2. Evaluating Creative Writing

Rubrics are one of the major tools for assessing writing which incorporate a set of prominent characteristics relevant to a specific type of discourse (Weigle, 2002). Vaezi and Rezaei (2019) developed a rubric for the evaluation of fiction writing fiction through nine elements, namely narrative voice, characterization, story, setting, mood and atmosphere, language and writing mechanics, dialogue, plot, and image. To ensure its validity, they further recruited a number of distinguished creative writing professors to review this assessment tool and comment on its appropriateness for measuring the intended construct. Biggs and Collis (1982) proposed to evaluate creative writing through the lens of the structural complexity of the product by utilizing the SOLO (Structure of the Observed Learning Outcome) taxonomy which buckets creative writing into incoherent (prestructural), linear(unistructural), conventional (multistructural), integrated (relational) and metaphoric (extended abstract). In prior work Rodríguez (2008) argued that narrative theory is key in teaching and grading creative writing. They emphasized how breaking down narrative elements such as plot, discourse-time, character, setting, narration, and filter delineates the tools authors use to effectively write fiction. In the creativity evaluation space Baer and McKool (2009); Amabile (1982) proposed The Consensual Assessment Technique as a method of assessing creative performance on a real-world task such as writing a poem or a story. Unlike tests from the Torrance Tests of Creative Thinking (TTCT), the CAT does not rely on any specific criteria or test scores. The method proposes that the most valid assessment of the creativity of an idea or creation in any field is the collective judgment of experts in that field. Unlike prior work, we aim to align the evaluation of creativity as a process to the evaluation of creativity as a product with feedback from experts building upon prior theoretical works such as the TTCT and the CAT. Unlike prior work from Vaezi and Rezaei (2019) where the rubric was created by surveying existing literature our work involved experts for creating the rubric from scratch without biasing their opinion or thought-process. Finally our rubric was validated by 5X more experts and across 3X more stories compared to that of Vaezi and Rezaei (2019).

3. Expert Evaluation of Language Model Generations

A cogent argument posits that the engagement with artistic prose isn’t confined solely to specialists, but non-experts also can be competent in assessing imaginative prowess. Inheriting from research standards for large-scale natural language processing tasks, the majority of studies assessing the quality of generations from LLMs evaluate model performance by collecting data from crowd workers (Nichols et al., 2020; Rashkin et al., 2020; Goldfarb-Tarrant et al., 2020; Roemmele and Gordon, 2018a; Yao et al., 2019). Nevertheless, it is critical to acknowledge that during narrative evaluation trials involving both teachers in English and participants from Amazon Mechanical Turk, the study by Karpinska et al. (2021) exhibited that AMT contributors, even when shortlisted via rigorous eligibility parameters (unlike teachers), struggle to discriminate between model generated text and human-crafted references. Clark et al. (2021) run a similar study assessing non-expert’s ability to distinguish between human and machine-authored text (GPT2 and GPT3) in three domains (stories, news articles, and recipes) and find that, without training, evaluators distinguished between GPT3 and human-authored text at random chance level. Mirowski et al. (2023) emphasized why crowd workers are not a good fit for evaluating AI-generated screenplays and instead engage 15 experts—theatre and film industry professionals—who have both experiences in using AI writing tools and who have worked in TV, film, or theatre in one of these capacities: writer, actor, director, or producer for evaluating LLM generated screenplays. Finally, a recent study by Veselovsky et al. (2023) highlighted the fact that approximately 33-46% of crowd workers on such platforms currently utilize large language models (LLMs) to complete any assigned task. Taking account of the above factors, for our work on evaluating short stories, we recruit creative writing experts ranging from professors to literary agents as well as MFA Fiction candidates.

Design Considerations

One of the main contributions of our work is the collection of 14 tests, referred to as the Torrance Test for Creative Writing (TTCW), to evaluate creativity in short fictional stories. These tests were formulated through collaboration with domain experts which we detail in Section 4, but in this Section, we first present the design principles that shaped the methodology and provide the desiderata for the tests, which are then empirically validated.

The 14 tests we propose are grounded in the Torrance Test for Creative Thinking (TTCT) (Torrance, 1966), which has been a cornerstone in the evaluation of creativity. TTCT provides measures to understand the creative process through tasks that encompass unusual uses of objects, scenarios, and out-of-the-box problems. TTCT is centered around evaluating four dimensions of creativity

Fluency. The sheer volume of meaningful ideas produced in reaction to a given stimulus.

Flexibility. The diversity of categories within the responses.

Originality. The uniqueness or novelty of answers.

Elaboration. The depth or granularity of details within the responses.

While the direct applications of TTCT might exhibit limitations in terms of applicability across diverse creative domains (Amabile, 1982; Baer and McKool, 2009), its fundamental dimensions have proven adaptable. Researchers have repurposed these dimensions effectively in diverse sectors like science education (Trisnayanti et al., 2019), content strategies in marketing (McIntyre et al., 2003), and even in human-computer interaction, particularly interface design (Bourdeau et al., 2020). In Section 4, we show how using creative writing experts we design domain-specific tests grounded in each TTCT dimension. Feedback from these specialists further underscores the pertinence of these dimensions in assessing creative writing.

A key consideration when designing tests to evaluate creativity is whether to center the evaluation on the cognitive process that leads to creativity or whether to evaluate the final artifact, which is a byproduct of the process (Mayers, 2007). Much prior work – including the TTCT – takes a design-centric approach, as it includes richer observation of the evaluated individual, which might not be captured in the final artifact. However, process-oriented evaluation is limited in several ways. First, process-oriented evaluation is inherently limited by the quality of the observation of the individual’s process. For instance, internal thoughts of the individual and other unrecorded activities can bias and lower the quality of the evaluation. Second, prior work has argued that neatly separating a process from an artifact is challenging, as the two are “tightly integrated” (Mayers, 2007), with “the creative process leaving traces within the artifact” (Murray, 2012). Finally, observing the process is not always possible, particularly when evaluating the creativity of a preexisting artifact (e.g., a short story written years ago), or evaluating black-box agents such as LLMs, whose process cannot be observed in an interpretable way. We, therefore, follow prior work Vaezi and Rezaei (2019); Rodríguez (2008) and design our creative writing evaluation to be artifact-centric.

In accordance with findings from prior work (Matell and Jacoby, 1971) which showed that reliability and validity are independent of the number of scale points used for Likert-type items, we stick to a binary scale. The evaluation within each Torrance dimension follows a similar procedure. Each dimension is associated with multiple binary questions, that represent individual tests. Each binary question is formulated as having a Yes/No answer such that an artifact receiving a “Yes” answer to a question corresponds to the artifact passing the test. Additionally, the Yes/No answer should be accompanied by a free-text rationale written by the evaluator which justifies the chosen binary label, with a length expectation of at least 1-3 sentences. For each test, the combination of a structured binary assessment and an open-ended rationale are complementary. The binary assessment can be used for quantitative assessment, such as measuring agreement amongst evaluators, or comparative evaluation of a story collection (such as the one we perform in Section 5.4), whereas the rationale can be used for qualitative assessment, such as understanding concrete reasons for the passing or failing of a test, such as the analysis we perform in Section 7.2 centering around the most common themes that lead to the passing or failing of a given test.

Each question is intended to be independent of other questions (i.e., no question is a prerequisite to another question), but the creative assessment of a given artifact requires completing all the TTCW. The final assessment of a given artifact is the number of tests passed by the artifact, with the general expectation that passing more tests is directly proportional to the creativity of the artifact. In other words, the passing or failing of any single test cannot be interpreted as a final assessment of the creativity of an artifact, but rather the number of tests passed can paint a more complete picture of the creativity of the artifact. Analysis in Section 5.4 confirms that our experts achieve on average moderate agreement on individual tests, but strong agreement when considering all tests in aggregate, confirming empirically the additive nature of the tests. In Section 4, we conduct a formative study with experts to formulate the fourteen TTCW that satisfy our design principles, which we then use in Section 5.4 to run a evaluation of short stories using TTCW with experts as assessors.

Formative Study: Formulating the Torrance Tests For Creative Writing

We restricted the involvement of participants in our formative study to only those possessing either a structured educational background in creative writing (for instance, a Master of Fine Arts in Creative Writing), traditionally published authors We do not recruit self-published authors, or lecturers/professors instructing Fiction Writing at the university level. We specifically chose this filtering criterion to restrict our selection pool to experts in the field thereby aligning with the Consensual Assessment Technique. Our recruitment resulted in participants who have published novels with leading publishing houses, students enrolled in top MFA programs in the United States, University professors teaching Fiction Writing, and screenwriters from prime-time networks. Participants were recruited through UserInterviews https://www.userinterviews.com, a professional freelancing website, and were paid $70 for taking part in the hour-long survey. Table 2 shows the background of the recruited participants. Our recruited participants span across different age groups, gender and professional expertise.

The formative study was structured in three parts. Initially, over video conferencing, experts were briefed on the study’s primary objective, which aimed at devising actionable metrics for assessing creative writing, emphasizing fiction. They were also introduced to the Torrance Tests and the four specific dimensions encompassed by it. In the subsequent phase, participants were emailed the URL to a web app where they were asked to input their measures in text. They were instructed to allocate a 20-minute window to articulate up to five distinct measures corresponding to each of the Torrance dimensions first. This task was conducted without a sample fiction to maintain abstraction. The final phase involved presenting the participants with a sample fiction piece https://www.newyorker.com/books/flash-fiction/the-mirror on the same web app for evaluation, retrieved from The New Yorker. Participants were encouraged to use this as a tangible example to refine and augment their initial measures, ensuring they were grounded and practical. As an outcome of our study, we received 126 measures from the participants across the four Torrance dimensions. The research was conducted at an institution that does not have an IRB approval process in place, but an Ethical Practices team reviewed the work and study protocols. We did not collect or share any PII during data collection, and participants could choose not to complete the survey and still receive a payment.

The measures derived from the participants exhibited a considerable degree of semantic congruence. For example, W2 proposed one way to measure Originality in Creative Writing as “Shows an innovative use of form/structure.” while W4 proposed a similar measure Formal or stylistic novelty. W6 proposed one way to measure Elaboration depending on whether “The story has developed 3D characters.” while W3 proposed an exactly similar measure but phrased it as “Does the piece make a flat character complex?”. To consolidate these measures and develop a framework of the underlying tests for measuring creativity, we use a general inductive approach for analyzing qualitative data (Thomas, 2006). Following this method, three authors independently read all of the measures and assigned each measure an initial potential low-level group. Then, through repeated discussion, we reduced category overlap and created shared low-level groups associated. Finally, these low-level groups were collected into high-level groups, and a name was proposed for each group that encapsulates a generalized representation of the measures within the group. During the later meetings, an American Novelist and Creative Writing Professor were present to give further insights into the data.

In total, the tagging process yielded 14 distinct groups, 5 in the Fluency dimension, and 3 in Flexibility, Originality, and Elaboration. Table 2 presents the name we assigned to each group and the study participants that proposed a measure tagged within this group. For every group, we had a list of expert-suggested questions speaking about the same artifact. To choose a representative measure for every group, we selected the most well-articulated measure (in terms of word count). If the measure was already suggested as a question, we keep it intact; otherwise, we used the GPT4 model to convert a measure to a Yes/No question. For example, Originality suggests that the piece isn’t cliche →\rightarrow Is the story an original piece of writing without any cliches? Next, we list each TTCW and provide the necessary background to contextualize the test. It should be noted that our tests are additive in nature. Failing a particular test should not be interpreted as the fact that the story doesn’t present any creative elements. Instead, it should be considered, by counting the number of distinct tests passed by a given story to obtain a more calibrated understanding of the creativity in a given piece.

2. The Torrance Test for Creative Writing

: Compared with both reading and speaking fluency, writing fluency has always been traditionally harder to define (Abdel Latif, 2013). Our 5 measures across this dimension each look at individual aspects of creative writing.

This measure refers to the manipulation of time in storytelling for dramatic effect. Essentially, it is about controlling the perceived speed and rhythm at which a story unfolds. A skilled writer can manipulate the relationship between these two to affect the pacing of the narrative, either speeding it up (compression) or slowing it down (stretching). This technique plays a crucial role in shaping the reader’s experience and engagement with the story. To assess narrative pacing, W4 suggested looking “Compression/stretching of time (story time vs. real world time)” while W6 and W7 advised to “control the speed at which a story unfolds https://www.writingclasses.com/toolbox/articles/stretching-and-shrinking-time.”

A ‘Scene’ is a moment in the story that is dramatized in real-time, often featuring character interaction, dialogue, and action, while ‘Exposition’, on the other hand, involves summarizing events or providing information like character history, setting details, or prior events. The right balance between scene and summary/exposition can vary depending on the story, but in general, it’s essential for maintaining a good pace, keeping the reader engaged, and delivering necessary information Burroway et al. (2019) https://creativenonfiction.org/syllabus/scene-summary/. W2 strongly felt that fluent writing needs to “display awareness and insight into the balance between scene and summary/exposition in the story.” while W5 and W8 emphasized the need for “Enough dialogue to compensate for backstory.”

Eminent novelist Milan Kundera said “Metaphors are not to be trifled with. A single metaphor can give birth to love.”. Sophisticated use of literary allusion or figurative language such as metaphor/idioms often add depth, interest, and nuanced meaning to any creative writing. It allows for a richer reading experience, where the literal events are imbued with deeper symbolic or thematic significance. W4 and W2 both emphasized the presence of “Sophisticated use of idiom, metaphor, and literary allusion or Surprising, skilled, and complex use of metaphor/simile/allusion as a way to measure Fluency in creative writing.

In her New Yorker essay “On Bad Endings” (Accocela, 2012) Accocela writes “Another possibility is that the author just gets tired. I review a lot of books, many of them non-fiction. Again and again, the last chapters are hasty and dull. ‘I’ve worked hard enough,’ the author seems to be saying. ‘My advance wasn’t much. I already have an idea for my next book. Get me out of here’.” If the writer ends the piece simply because they are “tired of writing”, the conclusion might feel abrupt, disjointed, or unfulfilling to the reader. This is one of the important factors of creative writing fluency. A strong ending offers a sense of closure, ties up the central conflicts or questions of the story, and generally leaves the reader feeling that the narrative journey was worthwhile and complete. For this measure of Fluency, W3 asked “Does the writer know how to end the piece not because they’re tired of writing, but because they have come to the moment the entire piece has been leading us towards?”

Narrative coherence is the degree to which a story makes sense https://en.wikipedia.org/wiki/Narrative_paradigm. A well-crafted story usually follows a logical path, where the events in the beginning set up the middle, which then logically leads to the end. Every scene, character action, and piece of dialogue should serve the story and propel it forward. Well-written stories have an underlying unity that binds the elements together. W3 strongly advocated for this measure by saying “Does the piece hold together? In other words, does the beginning lead through the middle to the end in a way that feels deliberate and intentional? This is the difference between several pages of writing and a PIECE of writing. Great writing errs on the side of unity over disorder.” while W7 suggested the importance of this measure through “a logical flow.”

Flexibility is often referred to as the ability to look at something from a different angle or point of view. In the context of creative writing, our participants agreed on 3 distinct measures of Flexibility.

An omniscient narrator is the all-knowing voice in a story that can convincingly and accurately depict a wide range of character viewpoints, including those of characters who may be morally ambiguous, difficult, or otherwise unappealing. As stated in Friedman (1955) an omniscient narrator enhances a sense of reliability or truth within literary works since readers are given deeper insights into many characters. The multiple viewpoints feel more objective because readers have access to multiple interpretations of events and can thus decide how they feel about each character’s perspective. W3 wanted “a writer to be able to inhabit various perspectives, even unlikable ones.” while W6 suggested “Flexibility of voice. Can the author inhabit the consciousness of different characters and not just likable ones?”

Emotional flexibility is asking whether the piece of writing effectively balances action and introspection, and if it portrays a broad and realistic spectrum of emotions as corroborated by W3. Exteriority refers to the observable actions, behaviors, or dialogue of a character, and the physical or visible aspects of the setting, plot, and conflicts.Interiority, on the other hand, pertains to the inner life of a character — their thoughts, feelings, memories, and subjective experiences. A balance between these two aspects is crucial in creating well-rounded characters and compelling narratives. As stated in Campe and Weber (2014) if a story is too heavy on exteriority, it may feel shallow or lack emotional depth. If it leans too much on interiority, it could become overly introspective and potentially lose the momentum of the plot.

2.8. Structural Flexibility(TTCW Flexibility3): Does the story contain turns that are both surprising and appropriate?

A good piece of creative writing often has plot twists, character developments, or thematic revelations that surprise the reader, subverting their expectations in a thrilling way. However, despite the surprises and twists, the turns in the story must also make sense within the established context of the story’s universe, its characters, and its themes. It shouldn’t feel like the writer has broken the rules they’ve set up, or made a character behave inconsistently without reason, simply for the sake of shock value. In order for any writing to be structurally flexible W3 wanted to ensure that “the writer capable of making turns in the work that are both surprising and appropriate while W1 required the presence of “New twist in the story that are believable to measure structural flexibility

Creative writing requires originality, or the ability to generate unique ideas (Ward et al., 1999). Our participants suggested three unique ways in which they look for originality in creative writing.

In his book “Literature and the Brain” well-known literary critic and scholar Norman Holland discusses how stories stimulate the mind and impact readers Holland (2009). A good story that offers a deeper understanding of human nature, cultural insights, unique viewpoints, or even the exploration of new ideas and themes has a lasting impact on its reader and society. In “Poetic Justice”, prominent philosophers Martha Nussbaum explores how the literary imagination is an essential ingredient of public discourse and a democratic society Nussbaum (1997). As such originality in theme and content is an important measure of creative writing. In the words of W6 originality in theme meant “New brilliant ideas about the future and humanity (mostly in speculative fiction). Does the writing have an original message? while W3 mentioned “Do I feel I am learning something new from the piece? What is the purpose of putting it into the world?”

2.10. Originality in Thought (TTCW Originality2): Is the story an original piece of writing without any cliches?

A cliche is an idea, expression, character, or plot that has been overused to the point of losing its original meaning or impact Fountain (2012). They often become predictable and uninteresting for the reader. In his book Clark (2008) eminent American writer, editor, and writing coach: Roy Peter Clark advised writers to strictly avoid cliches because they often indicate a lack of original thought or laziness in language use. Originality suggests that the piece isn’t cliche. Several experts agreed on this measure with W3 saying “Originality suggests that the piece isn’t cliche. while W4 required “Plot that is surprising rather than cliche”. W6 emphasized originality in thought through “writing that is not cliche, not recycled tropes, and does not contain stereotyped one-dimensional characters

2.11. TTCW Originality in Form & Structure (Originality3) : Does the story show originality in its form?

In his book Boardman (1992) Frederic Jameson highlighted the complexities of postmodern literature, where the blurring of genres and innovation in form was a key characteristic Jameson (1991). Originality in form has also been accomplished by the unconventional use of format, genre, or narrative structure or arc. For instance, the Pulitzer-winning book The Color Purple by Alice Walker is told through a series of letters written by the protagonist. Neil Gaiman’s American Gods on the other hand combines elements of fantasy, mystery, and mythic fiction in unexpected ways. The Sound and the Fury by William Faulkner deviates from the traditional plot structure by presenting a narrative that unfolds through the stream of consciousness of different characters. The goal of originality in form or structure is often to provide a fresh reader experience, challenge conventional reading expectations, or to create a deeper or more complex exploration of the story’s themes. This was corroborated by W2 and W5 who wanted to see “formal or stylistic novelty. W3 also wanted to see if A piece demonstrates original ways to work in sometimes tired forms (such as the short story)

Elaboration in creative writing is the process of adding details and information to a story to make it more interesting and engaging. This can be done by describing the setting, characters, and action in more detail, or by adding dialogue, thoughts, and feelings to the story.

American poet and memoirist Mark Doty discuss the importance of creating a vivid, immersive reality at the sensory level through the use of detailed, evocative description (Doty, 2014). An effective writer often uses sensory details to paint a detailed picture of the story’s environment, making it feel tangible and real to the reader. This level of detail contributes to the believability of the world, even if it is a completely fictional or fantastical setting. It helps the reader to suspend disbelief and become more deeply invested in the narrative. W2 suggested measuring elaboration through “Use of setting to inform the action and atmosphere of the story without succumbing to pathetic fallacy”. W6 further confirmed the same by saying “There is a setting. The short story is rooted in a time and a place. If it’s science fiction or fantasy there will be some world building.”. W5 further mentioned “the presence of a self-consistent world” as a measure of elaboration while W3 asked “How does the writer make the fictional world believable at the sensory level?

A ’flat character’ is typically a minor character who is not thoroughly developed or who does not undergo significant change or growth throughout the story. They often embody or represent a single trait or idea, and they’re only used to advance the plot or highlight certain qualities in other characters. A ’complex character’ on the other hand also known as a round character, has depth in feelings and passions, has a variety of traits of a real human being, and evolves over time. Forster (1927); Fishelov (1990); Currie (1990) highlights that any creative piece of fiction or non-fiction tends to be more engaging to the reader when authors can take a character who initially appears to be one-dimensional or stereotypical (flat) and add depth to them, as it mirrors the complexity of real people. Multiple experts emphasized on the importance of character development. For instance, W6 thought “The story should have developed 3D characters. while W2 and W8 asked for “Complex character development that avoids stereotype, generalization, trope, etc”. W3 wanted to see “Does the piece make a flat character complex?

In Ernest Hemingway’s short story “Hills Like White Elephants,” the couple’s conversation about seemingly unrelated topics implies a much deeper and more serious discussion about abortion. Their actual dialogue never directly addresses this issue, but it’s heavily suggested through what’s left unsaid — the subtext. Effective writing often operates on both surface and subtext levels. The surface text keeps the reader engaged with the plot and characters, while the subtext provides depth, complexity, and additional layers of interpretation, contributing to a richer and more rewarding reading experience Kochis (2007); Phelan (1996). W3 asked for rhetorical complexity by saying “Is there a sense of the piece being complex in such a way that it needed to be written as a whole, not simply paraphrased or summed up? while W4 mentioned that “Text should operate at multiple levels of meaning (surface and subtext).”

TTCW Implementation with Experts as Assessors

In this Section, we detail our implementation of creative writing evaluation using the TTCW framework derived previously. We first carefully select the 48-short stories included in the evaluation, then go over the 2.5-hour study design protocol that our 10 experiment participants followed, and finally analyze the results based on the 2,000+ individual tests administered by the experts.We leverage the collected annotations to answer the following research questions:

Are the human-written New Yorker stories more likely to pass individual TTCW than LLM-generated stories? If so, which tests demonstrate the most significant gaps?

Is the TTCW-based creative evaluation consistent and reproducible? In other words, is there agreement among expert annotators when they perform tests for similar stories?

Which LLMs perform better in TTCW evaluation, and are there specializations observed, with some different LLMs performing better on different Torrance dimensions?

Prior studies in the creativity of model-generated text have shown that technologies such as LLMs are capable of generative long and coherent stories (Yang et al., 2023, 2022), as well as act as assist and collaborate with creative writers (Yuan et al., 2022; Ippolito et al., 2022; Mirowski et al., 2023). For practical considerations, these studies limit their evaluation solely to model-generated content, typically only claiming relative improvements from one system to another, and do not establish whether a gap remains between high-quality human-written stories and model-generated text. Our research employs a rigorous evaluation protocol that juxtaposes human-written short stories against those constructed by LLMs to discern the potential gap, if any, in their creative quality. The dataset selection process is visually summarized in Figure 2. We first collect 12 short stories from The New Yorker collection of short stories https://www.newyorker.com/tag/short-stories. Our selection criterion involved choosing short stories with diverse authors and plots. Spanning from August 13, 2020, to May 8, 2023, these stories (ranging between 1000 to 2,400 words as delineated in Figure 5 in the Appendix) include compositions from acclaimed authors such as Haruki Murakami to Nobel laureate Annie Ernaux. The titles and a one-sentence plot summary (generated by GPT-4 and verified by humans) of included New Yorker short stories are listed in Table 10 in Appendix.

We prompt three top-performing LLMs: GPT3.5, GPT4, and Claude V1.3 to generate a story of similar length to each New Yorker story, based on the one-sentence plot summary. We note that the choice to condition model-generated stories on the plot of the New Yorker story is an important design consideration of our evaluation. First, LLMs are known to have limited ability in devising original plotlines, as highlighted in previous research (Ippolito et al., 2022). Second, it creates groups of stories that center around a common plot, allowing to dissociating evaluation of a story’s form (i.e., creative writing), from the evaluation of plot-line creativity.

Although LLMs were prompted to generate stories of a given length, initial experimentation by the authors of the paper revealed that the LLMs typically generate stories that can be 20-50% more concise than intended. Length is a known confounding factor in text generation evaluation, for example with work showing that evaluators systematically prefer longer summaries as they tend to be more informative (Stiennon et al., 2020). To address this limitation, we employed an iterative mechanism that prompted the LLM to iteratively expand on its initial story until the divergence in word count between the AI-generated and its paired human-written story was less than 200 ( See Prompt in Table 11 in Appendix). In our processing, all LLMs were able to converge to the desired length in at most 20 iterations. The procedure yields a total of 12 story groups, each consisting of one New Yorker story, and three LLM-generated stories, all following a common plot-line and having very similar length, for a total of 48 short stories. We also experimented with a few other choices in prompt design such as adding You are an expert of creative fiction writing to the beginning of the prompt or demonstrating an example New Yorker story in the prompt but these did not lead to any appreciable difference in the quality of the stories based on preliminary evaluation.

2. Evaluation Protocol

We want these tests described above to be understandable by both other creative writing experts or even LLMs, such that they can be used for evaluation purposes. An expert suggested questions for empirically evaluating creative writing might frequently elicit ambiguity in Large Language Models or even other creative writing experts. In order for LLMs or other experts to comprehend the suggested questions in Section 4.2, we attempt to expand them by adding more details. Recent pre-trained LLMs (e.g., GPT-4 (OpenAI, 2023) GPT3.5 (OpenAI, 2022)) can engage in fluent, multi-turn conversations out of the box, substantially lowering the data and programming-skill barriers to creating passable conversational user experiences. People can improve LLM outputs by prepending prompts—textual instructions and examples of their desired interactions—to LLM inputs. The prompts steer the model towards generating the desired outputs, raising the ceiling of what conversational UX is achievable for non-AI experts. To elucidate these questions we prompt GPT4 with the following instruction: What do creative experts mean when they say the following: {{expert question}}. Once GPT4 gives a response 3 domain experts carefully verify the response and edit it where required. Table 3 (Row2) shows the human-verified GPT4 expanded expert measure in response to the input prompt. Table 3 (Row3; Human Instruction) shows the final instruction given to human experts during the evaluation of our stories that contains the expanded expert measure {{M}} in addition to the original yes/no question. More examples of Human Instructions for the remaining 13 tests are provided in Section A.6 in the Appendix.

2.2. Expert Evaluation Protocol

We developed an evaluation protocol tailored specifically for experts in the domain of creative writing. The protocol, designed to be completed in approximately 2 to 2.5 hours, centered around a rigorous assessment of tuples of four distinct stories (one New Yorker story and the associated LLM-generated stories by the three LLMs) using the TTCW. The study was structured as follows:

The four stories in a group were shuffled, and anonymized (i.e., the author of the story was not visible to the evaluator).

The expert evaluator read the first story in its entirety and then administered the fourteen TTCW tests by assigning a Yes/No label and providing a justification for the label.

Upon completing the evaluation of the first story in the group, the evaluator proceeded to read and evaluate the second, third, and fourth stories in the group respectively. The evaluators were also allowed to edit their responses at any point of time during the entire process.

Once the evaluator had completed the TTCW evaluation of the four stories within a group, they were asked to rank all four stories in terms of subjective preference, and were asked to make an estimated guess of each story’s origin: choosing from “An experienced writer”, “An amateur writer”, or “Written by AI.” The exact formulation of each question is given in Figure 6 in the Appendix.

In preliminary trials conducted by the authors of the paper, the entire task completion was observed to range from 2-2.5 hours. Consequently, an $80 remuneration was determined to appropriately acknowledge the expertise of participants and encourage them to provide detailed justifications in their TTCW assessments. We note that participants were not provided details on the process used to create the groups of four stories, and were not told that each group consisted of one human-written story and 3 LLM-generated stories. Participants were explicitly instructed to avoid using search engines, which might reveal the origin of the New Yorker story. We also ensured beforehand that the participants were not familiar with any of the stories within a given group. Finally, participants were permitted to take breaks during the study but were encouraged to complete the entire task within a 24-hour window, so they would clearly remember each story when completing the final comparative task.

3. Participant Recruitment

To test the robustness and validity of TTCW-based evaluation, we chose to recruit a new set of experts to conduct the evaluation and not re-hire the ones from our formative study that played a role in the creation of the tests. We posit that such a choice demonstrates the fact that the TTCW can be administered by any knowledgeable expert provided solely with the tests and their expanded explanations. We recruited 10 participants on the User Interviews platform and listed their background in Table 4. Four of these participants are associated with the creative writing departments at leading American academic institutions, with considerable experience conducting undergraduate and graduate-level courses. Two participants function as literary agents at a top-tier, full-service US literary agency representing well-recognized authors, and four are professional writers with a Master of Fine Arts in Fiction or Poetry.

To experimentally analyze the reproducibility and validity of the TTCW, we randomly assigned each story group to 3 distinct experts. This allows us to study the agreement levels between experts on individual tests as well as in aggregate. After each task, an expert participant was given the option to receive another group of stories, and our participants completed on average 3.6 tasks over a period of 3 weeks, for a total of 36 assessments (i.e., 3 for each of the twelve groups). Because each assessment contains four stories, and each story was evaluated using the 14 TTCW, we collected a total of 2,016 binary labels and expert-written justifications for these labels.

4. Results

Table 5 summarizes the average passing rate on the 14 TTCW for each of the four story types (GPT3.5, GPT4, Claude-v1.3, and New Yorker). Passing rate here corresponds to the percentage of time expert participants answer ‘Yes’ to an individual TTCW for any given story. The New Yorker stories widely achieve the highest passing rate on all fourteen tests, with an overall pass rate of 84.7%. In other words, individual New Yorker stories are assessed to pass 11.9 of the 14 TTCW on average. When examining performance on individual tests, no test receives a pass rate of 100%, confirming that no test is an absolute requirement in high-quality creative writing, and experimentally justifying the need to conduct the TTCW tests as a set (Design Principle 4).

Moving to the performance of the LLM-generated stories, passing rates are much lower, with GPT3.5 stories passing less than 10% of TTCW, while GPT4 and Claude v1.3 are closer to 30.0%. In other words, LLM-generated stories pass between a third and a tenth of the TTCW compared to human-written New Yorker stories. When breaking down LLM-story performance across the Torrance dimensions, all models achieve their highest pass rate on the Fluency dimension, and Claude v1.3 achieves the highest performance on average across Fluency, Flexibility, and Elaboration, while GPT4 scores highest on the Originality dimension. This experimental finding is surprising as Claude v1.3 is an LLM that is smaller in size(52B) than GPT4 https://the-decoder.com/gpt-4-architecture-datasets-costs-and-more-leaked/.

4.2. Reproducibility of TTCW

Since we collected three independent TTCW evaluations for each story, we can report the agreement levels of experts when they conduct the tests individually, and their assessment in aggregate. We compute the Fleiss κ\kappa agreement across all annotations and report the interrater agreement level of each test in Table 5. Individual test agreement ranges from 0.27 to 0.66, and averages at 0.41, suggesting moderate agreement on most of the individual TTCW.

Since the TTCW are designed to be additive, we further compute an aggregate score for each story by counting the number of TTCW tests a story passes. We visualize the results of this aggregate measure in Figure 3. Since the aggregate measure is numerical (ranges from 0 to 14), we use Pearson correlation to measure agreement among experts. At this aggregate level, we obtain a correlation(ρ\rho) of 0.69, showing strong agreement among experts on the number of tests a story passes. In other words, even though experts reach slightly lower agreement on which exact TTCW a story passes or fails, they achieve strong agreement on the number of tests a story passes overall. This experimental finding confirms the importance of Design Principle 4, and the need for the tests to be performed as a set to achieve a reproducible evaluation of creativity for a given short story. When the objective is to evaluate the broad creativity in a short story, it is our recommendation that all fourteen tests should be administered as a set by one expert annotator, rather than by different experts or administering only an individual test, as this increases reproducibility of results.

4.3. Comparative Evaluation Results

The final portion of the evaluation protocol asks expert participants to rank the four shuffled stories in terms of subjective preference, as well as guess each story’s origin between “An experienced writer”, “An amateur writer”, or “Written by AI”. Figure 4 summarizes the results from this final portion of the study.

Looking at the ranking results, human-written New Yorker stories were ranked as the most preferred story 89% of the time, while the GPT-3.5-generated stories rank as least preferred roughly two-thirds of the time. When comparing between GPT-4 and Claude, Claude is almost twice as likely to rank as second (behind the human-written story) and was the most preferred on three of the four assessments in which the New Yorker story was not chosen as the most favored. These ranking results confirm and accentuate the observation from the test passing rates analysis that Claude V1.3 generates higher-quality short stories than models in the GPT family.

The attribution results paint a similar picture, with New Yorker stories predominantly attributed to an experienced writer, while LLM-generated stories get attributed to AI or an amateur writer. Interestingly, Claude V1.3 is more likely to be attributed to an amateur writer than an AI, whereas GPT3.5 and GPT4 stories are 80%+ attributed to AI. One hypothesis for such a behavior could be that the participants in our study might are more familiar with text written by OpenAI models, as these models are commercially more successful, providing an element of surprise to Claude-written text.

TTCW Implementation with LLMs as Assessors

Expert annotation such as the one we perform in Section 5 is costly: based on our evaluation protocol, evaluating a 1500-2500 word short story with a qualified expert costs $20, and requires roughly 30 minutes of the expert’s time. Prior work has shown the promise of using LLMs in text evaluation. For instance, GPT3.5 and GPT4 are effective at evaluating the factual consistency of a summary to its document (Laban et al., 2023), or measuring the coherence of a summary (Gao et al., 2023). GPTEval (Liu et al., 2023) employs the framework of using large language models with chain-of-thoughts (CoT) (Wei et al., 2022) to assess the quality of NLG outputs. Recent work has also applied LLM-based evaluation to the creative domain (Rajani et al., 2023), claiming that GPT4 can achieve a high correlation with humans when evaluating brainstorming or creative generation tasks. In this section we describe our TTCW implementation with LLMs as assessor to understand LLMs ability to assess creative writing.

We apply a similar evaluation protocol to the one described in Section 5.We use the same data selection and the same three LLMs: GPT3.5, GPT4, and Claude, and prompt them to answer the 14 individual TTCW tests for the 48 stories in our collection.

Prior work Wei et al. (2022) has shown how generating a chain of thought – a series of intermediate reasoning steps – significantly improves the ability of large language models to perform complex reasoning. Taking advantage of this we design the prompts/instructions for large language models in a slightly different fashion than for the human experts as can be seen in Table 3(Row 3 (Human Instruction) vs Row 4 LLM Instruction). To help the model make an informed decision we first ask it to list out elements specific to any given test such as “elements in the story that call to each of the five senses” for the World Building and setting test followed by asking it to decide overall and then provide its reasoning before choosing an answer between ‘Yes’ or ‘No’. The exact prompt contains (1) the story, (2) the expanded TTCW context, (3) the TTCW question, and (4) a LLM-specific instruction guiding the model to perform the task in a chain-of-though manner. We then measure model agreement with the majority vote of the three experts that conducted the test, using Cohen’s Kappa.

Table 6 summarizes correlation results between LLM-based and expert-based TTCW assessments. On average, we find that none of the LLMs produce assessments that correlate positively with expert assessments, with correlation averages close to zero. GPT4 is the only model to obtain correlations above 0.2 on two of the fourteen tests, yet this still does not qualify as moderate agreement.This empirical result contrasts with prior work: even though GPT4 has been shown to have some ability to evaluate creativity in short-form tasks (such as responses with less than 100 words) (Rajani et al., 2023), our work shows that this result does not extend to longer-form evaluation. We note that although the prompts we utilized were zero-shot in nature (i.e., these prompts did not include example binary labels and justifications from experts), we experimented with few-shot prompts for a couple of the TTCW tests and did not obtain any significant correlation gains.

Yet LLM-administered TTCW would be a crucial building block in improving model-generated stories. Assuming that an automated method could produce reliable TTCW outcomes, it could be used in iterative algorithms such as Self-Refine (Madaan et al., 2023) to iteratively edit a draft story until it passes a large proportion of tests. With this in mind, we release the TTCW benchmark which contains all binary judgments we collected and expert justifications, with the hope that the community can use it as a tool to track progress in the evaluation of the creative capabilities of LLMs.

To get a deeper understanding of how expert and LLM explanations differ we take a closer look at them. We explicitly prompted LLMs to do step-by-step reasoning before arriving at any verdict and this was often reflected in the explanations. The LLM-generated explanations were procedural and typically lengthier than expert explanations. While there were not any specific instructions given to experts about the length of the explanations we asked them to provide necessary details justifying their decision. Table 13 in the Appendix shows such an example.

Discussion

In an optional exchange with the expert participants who participated in the annotation of Section 5, participants were given the opportunity to describe how they differentiated between AI-generated and human-written stories. Table 14 in the Appendix lists the replies of five experts, which we color-coded to highlight the recurrent issues that led them to believe that a story is written by an AI. In particular, E5, E4, E1 thought AI struggles at Narrative Ending. E5 and E4 highlighted that AI-written stories would forestall the ending by getting bigger in scope. E1 highlighted that AI-written stories would have multiple disparate endings. E5, E4, E1, and E2 all highlighted that AI-written stories would often contain abstruse and incoherent metaphors that do not add meaning or extremely cliched or simple metaphors thereby demonstrating poor Language Proficiency and use of Literary Devices. E1 highlighted one such example in a story - However, she managed to laugh louder and louder until her laughter transformed into an embrace of the sun’s atmosphere.

E1 and E2 further highlighted that characters in AI-written stories have poor Rhetorical Complexity and are often lacking in subtext. E1 further added that AI-written stories operate in a nearly opposite and Drax-like fashion in which there is only literal meaning, and that literal meaning is often nonsensical, or at least presented without any of the context that might make it seem like something a human would say or do. E1 highlighted an exchange below from an AI-written story

Sarah: “We’ve been avoiding the inevitable, Max. During our time here we’ve grown closer, and now that it’s almost over, we can’t just pretend like it never happened.”

Max: “I understand, Sarah. But how do we move forward? How do we navigate this complexity without unraveling everything we’ve built?”

where he exclaimed that these statements make hackish sense as clumsy exposition directed at the reader, and no sense at all as sentences spoken from one alleged human being to another. Both E5, E3, and E1 agreed on poor Character Development in AI-generated stories where a character would appear and then disappear without having any impact. E3 and E2 also highlighted issues in Narrative Pacing where stories would either spiral into a repetitive pattern or rapidly accelerate through time after the first scene. E1 and E4 also highlighted Unusual Syntax in sentence structure in AI-generated stories and repetition of certain words and phrases across stories.

2. What can we infer from expert explanations of administering the TTCW?

Our annotation effort in Section 5 required experts to not only annotate for binary labels but also provide a justification paragraph accounting for their assessment. In Table 7 we provide examples of such justification for one of the TTCW tests for Originality. To gain insights into the main justifications experts provide for a story to pass or fail a TTCW test, we performed a manual thematic analysis of the expert justifications. We organized the results into a set of minimal phrases that often appear when a story passes or fails each TTCW test. Recent work has shown the utility of LLMs in clustering (Viswanathan et al., 2023). Based on these findings we asked the GPT4 model to cluster explanations across a given TTCW dimension into recurrent and broader representative themes. Three authors of the paper then manually verified these themes to ensure correctness. The outcome is summarized in Table 8 for Fluency and Flexibility tests, and Table 9 for Originality and Elaboration tests. The underlying themes found across the explanations reaffirm prior findings from Ippolito et al. (2022) where writers found LLM written stories experiencing difficulty to maintain a style/voice and easily reverting to tropes and repetition as well as those from Mirowski et al. (2023) where screenwriters complained about the lack of subtext and character motivation.

3. Towards Interactive LLM-based Creative Co-Writing

Our experimental results and the analysis of the expert explanations highlight the limitations of current LLMs in generating both high-quality fictional stories as well as assessing the creativity of such existing stories. With the rapid progression in LLM development, we make available a corpus containing expert evaluations of TTCW assessments. We believe such a contribution will facilitate the evaluation of the upcoming model’s abilities for creative writing assessment. If LLMs are capable of producing TTCW assessments that correlate positively with expert judgments, future work can explore new opportunities for creative LLM-based co-writing interfaces.

In particular, we envision LLMs providing assistance during Planning and Reviewing, crucial phases in the cognitive process theory of writing (Flower and Hayes, 1981). In prior work Gero et al. (2023) states that writers expressed the importance of specificity in the feedback instead of generic feedback like this might be a bit boring. We hope that the TTCW tests can provide the structure in future work looking to provide targeted feedback on writing. Further Ippolito et al. (2022) recently pointed out that professional writers constantly felt that LLM-generated text is rife with cliches and overused tropes. Metrics that quantify elements like originality in theme, structural flexibility, or rhetorical complexity could guide creative writing support toolshttps://www.sudowrite.com/ built using current LLMs, thereby improving planning and translation (Chung et al., 2021).

Limitations and Future Work

In our experimental design, we employed the default generation parameters for the models, specifically a temperature setting, (T=1.0T=1.0), aiming to evaluate the model’s capabilities in a non-optimized setting. Notably, variations in parameters such as temperature have been documented to potentially enhance the originality of generated content (Roemmele and Gordon, 2018b). Consequently, there exists a possibility that utilizing alternative generation parameters might lead to outputs that surpass a greater fraction of the TTCW. As the associated costs of conducting TTCW evaluations decrease, subsequent research endeavors could provide a more comprehensive insight into the influence of generation parameters on creativity within the context of the TTCW framework.

When selecting which expert measures would yield a TTCW test, we attempted to maximize coverage while minimizing overlap between the tests. Yet it is very likely that the passing of a test is correlated with another test, as the tests touch on common story elements (e.g., characters, prose). We did not investigate the degree of overlap between pairs of tests. We encourage future work to further refine the TTCW tests and propose additional tests.

We open-source the prompts used to instruct the LLMs to generate the stories as well as prompts to administer the TTCW tests, which we iterated on and reviewed carefully. While there is no upper bound on engineering the ‘best’ prompt for a particular task, the output of an LLM is still dependent on the input prompt (Zhou et al., 2022). The LLMs that we evaluate in the study are closed-source models and the quality of their generation has been observed to change over time (Chen et al., 2023a). We hope future work can explore further refining of the prompts and parameters used in our work, and exploring their impact on the results.

Our tests were explicitly designed for short fiction and we do not study the generalization of TTCW to other forms of creative writing. Empirical evaluation would need to be conducted to verify whether the TTCW is adequate and comprehensive to evaluate other forms of creative writing, including scripts, novels, or marketing material such as slogans. We posit there are specific metrics for each specific type of creative writing, which our current tests do not cover. For example, evaluating or critiquing poetry might require different fine-grained evaluation metrics compared to short stories.

Creativity emerges over time in a complex interplay of factors. Human judgments of creativity are often biased by personal tastes, expectations, and hindsight. The definition of an ‘expert’ or ‘amateur’ creative writer is not clear-cut in a field that has unclear professional delineations (Gero et al., 2023). Many successful writers retain full-time jobs as teachers, editors, or in unrelated professions, as few are able to make a living from their writing alone. Although we emphasize the validity of our results by computing agreement levels among recruited experts, we note that work in creativity should not solely aim to maximize agreement, since subjective differences are an inherent property of divergent thinking, which is central to creativity. Future work looking to expand our line of work should seek to select experts with diverse backgrounds, offering a more comprehensive view of the creative process.

Conclusion

In this work, we utilize the widely accepted Torrance Test of Creative Thinking originally designed for the evaluation of creativity as a process, and align it towards the evaluation of creativity as a product. In particular, we focus on short fiction writing and formulate the Torrance Test for Creative Writing (TTCW). Our collaborative process, involving creative writing experts in both the development and validation stages of the TTCW, has established the test as a robust instrument for assessing creativity in fictional short stories. The experimental data derived from the evaluation of both human-authored and LLM-generated stories offers a rich comparative analysis. It underscores the proficiency of seasoned writers in evoking creativity, outperforming LLMs by a considerable margin. Our analysis also highlights the disparity in creative prowess among different LLMs, revealing that while certain LLMs might exhibit proficiency in some dimensions of creativity, there exists a wide chasm between them and human expertise when assessed holistically. Notably, the TTCW tests also provided a granular perspective on the areas where LLMs falter the most in terms of creativity. Our secondary investigation into the feasibility of LLMs in reproducing expert assessments yielded that, at least in their current state, LLMs are not yet adept at administering TTCW tests. This indicates a dual challenge: not only are LLMs lagging in producing inherently creative content, but they also lack the finesse to evaluate creativity as experts do. We hope that our evaluation framework using TTCW, and findings on the creative capabilities of LLMs coupled with our dataset will steer future work and innovations in creativity research.

References

Appendix A Appendix

Table 10 shows the data used for conducting our evaluation. The 12 stories shown are taken from The New Yorker and summarized into single-sentence plots. These stories come from highly established literary experts acting as an upper bound for what it means to be creative. These stories span complex themes

A.2. Expert Perception on the TTCW tests

Since the experts listed in Table 4 were not involved in designing the rubric but evaluated several stories based on the rubric we asked them their overall thought about the rubric and any potentially crucial test we missed out on that they use to discriminate between good and bad writing.As can be seen in Table 12 in Appendix overall almost every expert agreed on the thorough and effective nature of our rubric. Many of them agreed on the fact that our rubric helped them to think about different aspects of storytelling in a more structured way. One of the difficult things about coming up with a rubric for creativity is ensuring coverage. Even though our rubric covers most aspects of creative writing, some experts such as E1 and E4 emphasized on the utility of Consistency of Voice and Diction as a measurable test. In E4’s words “Inconsistent voice and diction are sometimes/often notable in stories that aren’t very good, and when voice & diction are used beautifully, it enhances a story considerably”. E1 similarly exclaimed “One of the most meaningful aspects of high-quality literary writing is voice, which conveys qualities of proficiency, artistry, personality, and identity.”. We hope future work can adapt this as a meaningful test in addition to the tests covered in our rubric. Finally, some of the tests from our rubric can have potential overlaps as pointed out by E2. This is further corroborated by the similar numbers for Narrative Pacing and Scenes vs Exposition suggesting a strong correlation between the two.

A.3. Example LLM-generated and expert-written explanations for a TTCW assessment

In Table 13, we show examples of explanations that experts wrote in conjunction with a binary TTCW assessment they made on a story, as well as the corresponding LLM-generated explanations.

A.4. Expert Distinguishability Feedback

In Table 14, we color-coded the responses of five experts who participated in our large-scale TTCW-based annotation, to understand what were the key characteristics that differentiate AI-generated and human-written stories, in their opinion.

A.5. Can non-experts administer TTCW tests?

Recruiting experts for data annotation purposes is challenging, and costly, and must consider the time constraint put on the experts. Prior work has shown the potential of crowd-sourcing (through platforms such as Amazon Mechanical Turk) and the ability of non-experts to accomplish complex tasks as a crowd (Kittur et al., 2013), when following an appropriate workflow that iterates and validates the work on individual non-experts. Some prior work has even shown the validity of crowd-based feedback for writing tasks (Bernstein et al., 2010; Nebeling et al., 2016).

In this work, we chose to rely on experts for annotation, to maximize the validity of our experiments, and confirm whether experts with domain knowledge would reach satisfying agreement levels when evaluating stories with TTCW. Future work can leverage our open-sourced annotations to explore whether non-experts correlate with experts when performing TTCW evaluation, which could lead to more cost-effective TTCW evaluation.

A.6. Prompts for TTCW

All the instructions shown to creative writing experts and LLMs are given in the tables below.