Instructions for EMNLP 2023 Proceedings

Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4

Kent K. Chang
University of California, Berkeley
kentkchang@berkeley.edu
Mackenzie Cramer
University of California, Berkeley
mackenzie.hanh@berkeley.edu
Sandeep Soni
Emory University
sandeep.soni@emory.edu
David Bamman
1
University of California, Berkeley
dbamman@berkeley.edu


Abstract

In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query. We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with which passages of those books appear on the web. The ability of these models to memorize an unknown set of books complicates assessments of measurement validity for cultural analytics by contaminating test data; we show that models perform much better on memorized books than on non-memorized books for downstream tasks. We argue that this supports a case for open models whose training data is known.

1 Introduction↩︎

Research in cultural analytics at the intersection of NLP and narrative is often focused on developing algorithmic devices to measure some phenomenon of interest in literary texts[1][4]. The rise of large-pretrained language models such as ChatGPT and GPT-4 has the potential to radically transform this space by both reducing the need for large-scale training data for new tasks and lowering the technical barrier to entry[5].

None

Figure 1: Name cloze examples. GPT-4 answers both of these correctly..

At the same time, however, these models also present a challenge for establishing the validity of results, since few details are known about the data used to train them. As others have shown, the accuracy of such models is strongly dependent on the frequency with which a model has seen information in the training data, calling into question their ability to generalize[6][8]; in addition, this phenomenon is exacerbated for larger models[9], [10]. Knowing what books a model has been trained on is critical to assess such sources of bias[11], which can impact the validity of results in cultural analytics: if evaluation datasets contain memorized books, they provide a false measure of future performance on non-memorized books; without knowing what books a model has been trained on, we are unable to construct evaluation benchmarks that can be sure to exclude them.

In this work, we carry out a data archaeology to infer books that are known to several of these large language models. This archaeology is a membership inference query[12] in which we probe the degree of exact memorization [13] for a sample of passages from 571 works of fiction published between 1749–2020. This difficult name cloze task, illustrated in figure 1, has 0% human baseline performance.

This archaeology allows us to uncover a number of findings about the books known to OpenAI models which can impact downstream work in cultural analytics:

  1. OpenAI models, and GPT-4 in particular, have memorized a wide collection of in-copyright books.

  2. While others have shown that LLMs are able to reproduce some popular works[14], we measure the systematic biases in what books OpenAI models have seen and memorized, with the most strongly memorized books including science fiction/fantasy novels, popular works in the public domain, and bestsellers.

  3. This bias aligns with that present in the general web, as reflected in search results from Google, Bing and C4. This confirms prior findings that duplication encourages memorization[15], and provides a rough diagnostic for assessing knowledge about a book.

  4. Disparity in memorization leads to disparity in downstream tasks. GPT models perform better on memorized books than non-memorized books at predicting the year of first publication for a work and the duration of narrative time for a passage, and are more likely to generate character names from books it has seen.

While our work is focused on ChatGPT and GPT-4, we also uncover surprising findings about BERT as well: BookCorpus [16], one of BERT’s training sources, contains in-copyright materials by published authors, including E.L. James’ Fifty Shades of Grey, Diana Gabaldon’s Outlander and Dan Brown’s The Lost Symbol, and BERT has memorized this material as well.

As researchers in cultural analytics are poised to use ChatGPT and GPT-4 for the empirical analysis of literature, our work both sheds light on the underlying knowledge in these models while also illustrating the threats to validity in using them.

None

Figure 2: Sample name cloze prompt..

2 Related Work↩︎

2.0.0.1 Knowledge production and critical digital humanities.

The archaeology of data that this work presents can be situated in the tradition of tool critiques in critical digital humanities [17][21], where we critically approach the closedness and opacity of LLMs, which can pose significant issues if they are used to reason about literature [22][26]. In this light, our archaeology of books shares the Foucauldian impulse to “find a common structure that underlies both the scientific knowledge and the institutional practices of an age” ([27, p. 79]; [25]). to the inference tasks that involve them. Investigating data membership helps us reflect on how to best use LLMs for large-scale cultural analysis and evaluate the validity of its findings.

2.0.0.2 LLMs for cultural analysis.

While they are more often the object of critique, LLMs are gaining prominence in large-scale cultural analysis as part of the methodology. For example, some use GPT models to tackle problems related to interpretation, that of real events [28], or of character roles (hero, villain, and victim, [29]). Others leverage LLMs on classic NLP tasks that can be used to shed light on literary and cultural phenomena, or otherwise maintain an interest in those humanistic domains [30]. This work shows what tasks and what types of questions LLMs are more suitable to answer than others through our archaeology of data.

2.0.0.3 Documenting training data.

Training on large text corpora such as BookCorpus [16], C4 [31] and the Pile [32] has been instrumental in extending the capability of large language models. Yet, in contrast to smaller datasets, these large corpora are less carefully curated. Besides a few attempts at documenting large datasets [33] or laying out diagnostic techniques for their quality [34], these corpora and their use in large models is less understood. Our focus on books in this work is an attempt to empirically map the information about books that is present in these models.

2.0.0.4 Memorization.

Large language models have shown impressive zero-shot or few-shot ability but they also suffer from memorization [35], [36]. While memorization is shown in some cases to improve generalization [37], it has generally been shown to have negative consequences, including security and privacy risks [15], [38], [39]. Studies have quantified the level of memorization in large language models [9], [40] and have highlighted the role of both verbatim [41], [42] and subtler forms of memorization [43]. [14] in particular examines the legal issues surrounding copyright in foundation models, focusing in particular on the length of content extraction of copyrighted materials (including books as a case study) that poses the most risk to fair use determinations. Our analysis and experimental findings on the disparate sources of memorization and impact on research in cultural analytics add to this scholarship.

2.0.0.5 Data contamination.

A related issue noted upon critical scrutiny of uncurated large corpora is data contamination [44] raising questions about the zero-shot capability of these models [45] and worries about security and privacy [15], [38]. For example, [33] find that text from NLP evaluation datasets is present in C4, attributing performance gain partly to the train-test leakage. [41] show that C4 contains repeated long running sentences and near duplicates whose removal mitigates data contamination; similarly, deduplication has also been shown to alleviate the privacy risks in models trained on large corpora [46].

3 Task↩︎

3.1 Name cloze↩︎

We formulate our task as a cloze: given some context, predict a single token that fills in a mask. To account for different texts being more predictable than others, we focus on a hard setting of predicting the identity of a single name in a passage of 40–60 tokens that contains no other named entities. Figure 1 illustrates two such examples.

In the following, we refer to this task as a name cloze, a version of exact memorization[13]. Unlike other cloze tasks that focus on entity prediction for question answering/reading comprehension[47], [48], no names at all appear in the context to inform the cloze fill. In the absence of information about each particular book, this name should be nearly impossible to predict from the context alone; it requires knowledge not of English, but rather about the work in question.

The name cloze setup is also particularly suitable for our data archaeology for ChatGPT/GPT-4, as opposed to perplexity-based measures [49], since OpenAI does not at the time of writing share word probabilities through their API.

3.2 Evaluation↩︎

We construct this evaluation set by running BookNLP2 over the dataset described below in §4, extracting all passages between 40 and 60 tokens with a single proper person entity and no other named entities. Each passage contains complete sentences, and does not cross sentence boundaries. We randomly sample 100 such passages per book, and exclude any books from our analyses with fewer than 100 such passages.

We pass each passage through the prompt listed in figure 2, which is designed to elicit a single word, proper name response wrapped in XML tags; two short input/output examples are provided to illustrate the expected structure of the response.

In establishing baselines using the same evaluation set, predicting the most frequent name in the dataset (“Mary”) yields an accuracy of 0.6%. To assess human performance on this task, one of the authors of this paper followed the same instructions given ChatGPT/GPT-4, only using information present in the passage to guess the masked name (i.e., without using any external sources of information such as Google); the resulting accuracy is 0%, which suggests there is little information within the context that would provide a signal for what the true name should be.

Table 1: Top 20 books by GPT-4 name cloze accuracy.
GPT-4 ChatGPT BERT Date Author Title
0.98 0.82 0.00 1865 Lewis Carroll Alice’s Adventures in Wonderland
0.76 0.43 0.00 1997 J.K. Rowling Harry Potter and the Sorcerer’s Stone
0.74 0.29 0.00 1850 Nathaniel Hawthorne The Scarlet Letter
0.72 0.11 0.00 1892 Arthur Conan Doyle The Adventures of Sherlock Holmes
0.70 0.10 0.00 1815 Jane Austen Emma
0.65 0.19 0.00 1823 Mary W. Shelley Frankenstein
0.62 0.13 0.00 1813 Jane Austen Pride and Prejudice
0.61 0.35 0.00 1884 Mark Twain Adventures of Huckleberry Finn
0.61 0.30 0.00 1853 Herman Melville Bartleby, the Scrivener
0.61 0.08 0.00 1897 Bram Stoker Dracula
0.61 0.18 0.00 1838 Charles Dickens Oliver Twist
0.59 0.13 0.00 1902 Arthur Conan Doyle The Hound of the Baskervilles
0.59 0.22 0.00 1851 Herman Melville Moby Dick; Or, The Whale
0.58 0.35 0.00 1876 Mark Twain The Adventures of Tom Sawyer
0.57 0.30 0.00 1949 George Orwell 1984
0.54 0.10 0.00 1908 L. M. Montgomery Anne of Green Gables
0.51 0.20 0.01 1954 J.R.R. Tolkien The Fellowship of the Ring
0.49 0.16 0.13 2012 E.L. James Fifty Shades of Grey
0.49 0.24 0.01 1911 Frances H. Burnett The Secret Garden
0.49 0.12 0.00 1883 Robert L. Stevenson Treasure Island
0.49 0.16 0.00 1847 Charlotte Brontë Jane Eyre: An Autobiography
0.49 0.22 0.00 1903 Jack London The Call of the Wild

4 Data↩︎

We evaluate 5 sources of English-language fiction:

  • 91 novels from LitBank, published before 1923.

  • 90 Pulitzer prize nominees from 1924–2020.

  • 95 Bestsellers from the NY Times and Publishers Weekly from 1924–2020.

  • 101 novels written by Black authors, either from the Black Book Interactive Project3 or Black Caucus American Library Association award winners from 1928–2018.

  • 95 works of Global Anglophone fiction (outside the U.S. and U.K.) from 1935–2020.

  • 99 works of genre fiction, containing science fiction/fantasy, horror, mystery/crime, romance and action/spy novels from 1928–2017.

Pre-1923 LitBank texts are born digital on Project Gutenberg and are in the public domain in the United States; all other sources were created by purchasing physical books, scanning them and OCR’ing them with Abbyy FineReader. As of the time of writing, books published after 1928 are generally in copyright in the U.S.

5 Results↩︎

5.1 ChatGPT/GPT-4↩︎

We pass all passages with the same prompt through both ChatGPT (gpt-3.5-turbo) and GPT-4 (gpt-4), using the OpenAI API. The total cost of this experiment with current OpenAI pricing (\(\$0.002\)/thousand tokens for ChatGPT; \(\$0.03\)/thousand tokens for GPT-4) is approximately \(\$400\). We measure the name cloze accuracy for a book as the fraction of 100 samples from it where the model being tested predicts the masked name correctly.

Table 1 presents the top 20 books with the highest GPT-4 name cloze accuracy. While works in the public domain dominate this list, table 6 in the Appendix presents the same for books published after 1928.4 Of particular interest in this list is the dominance of science fiction and fantasy works, including Harry Potter, 1984, Lord of the Rings, Hunger Games, Hitchhiker’s Guide to the Galaxy, Fahrenheit 451, A Game of Thrones, and Dune—12 of the top 20 most memorized books in copyright fall in this category. Table 2 explores this in more detail by aggregating the performance by the top-level categories described above, including the specific genre for our subset of genre fiction.

GPT-4 and ChatGPT are widely knowledgeable about texts in the public domain (included in pre-1923 LitBank); it knows little about works of Global Anglophone texts, works in the Black Book Interactive Project and Black Caucus American Library Association award winners.

Table 2: Name cloze performance by book category.
Source GPT-4 ChatGPT
pre-1923 LitBank 0.244 0.072
Genre: SF/Fantasy 0.235 0.108
Genre: Horror 0.054 0.028
Bestsellers 0.033 0.016
Genre: Action/Spy 0.032 0.007
Genre: Mystery/Crime 0.029 0.014
Genre: Romance 0.029 0.011
Pulitzer 0.026 0.011
Global 0.020 0.009
BBIP/BCALA 0.017 0.011

5.2 BERT↩︎

For comparison, we also generate predictions for the masked token using BERT (only passing the passage through the model and not the prefaced instructions) to provide a baseline for how often a model would guess the correct name when simply functioning as a language model (unconstrained to generate proper names). As table 1 illustrates, BERT’s performance is near 0 for all books—except for Fifty Shades of Grey, for which it guesses the correct name 13% of the time, suggesting that this book was known to BERT during training. [50] notes that BERT was trained on Wikipedia and the BookCorpus, which [16] describe as “free books written by yet unpublished authors.”5 Manual inspection of the BookCorpus hosted by huggingface6 confirms that Fifty Shades of Grey is present within it, along with several other published works, including Diana Gabaldon’s Outlander and Dan Brown’s The Lost Symbol.

6 Analysis↩︎

6.1 Error analysis↩︎

We analyze examples on which ChatGPT and GPT-4 make errors to assess the impact of memorization. Specifically, we test the following question: when a model makes a name cloze error, is it more likely to offer a name from a memorized book than a non-memorized one?

To test this, we construct sets of seen (\(S\)) and unseen (\(U\)) character names by the models. To do this, we divide all books into three categories: \(M\) as books the model has memorized (top decile by GPT-4 name cloze accuracy), \(\neg M\) as books the model has not memorized (bottom decile), and \(H\) as books held out to test the hypothesis. We identify the true masked names that are most associated with the books in \(M\) — by calculating the positive pointwise mutual information between a name and book pair — to obtain set \(S\), and the masked names most associated with books in \(\neg M\) to obtain set \(U\). We also ensure that \(S\) and \(U\) are of the same size and have no overlap. Next, we calculate the observed statistic as the log-odds ratio on examples from \(H\): \[\begin{align} o &= \log \left(\frac{Pr \left(\hat{c} \in S | error\right)}{Pr \left(\hat{c} \in U | error\right)}\right), \end{align}\] where \(\hat{c}\) is the predicted character. To test for statistical significance, we perform a randomization test [51], where the observed statistic \(o\) is compared to a distribution of the same statistic calculated by randomly shuffling the names between \(S\) and \(U\).

We find that for both ChatGPT (\(o=1.34, p < 0.0001\)) and GPT-4 (\(o=1.37, p < 0.0001\)), the null hypothesis can be rejected, indicating that both models are more likely to predict a character name from a book they had memorized than a character from a book they have not. This has important consequences: these models do not simply perform better on a set of memorized books, but the information from those books bleeds out into other narrative contexts.

6.2 Extrinsic analysis↩︎

Why do the GPT models know about some books more than others? As others have shown, duplicated content in training is strongly correlated with memorization[15] and foundation models are able to reproduce popular texts more so than random ones[14]. While the training data for ChatGPT and GPT-4 is unknown, it likely involves data scraped from the web, as with prior models’ use of WebText and C4. To what degree is a model’s performance for a book in our name cloze task correlated with the number of copies of that book on the open web? We assess this using four sources: Google and Bing search engine results, C4, and the Pile.

For each book in our dataset, we sample 10 passages at random from our evaluation set and select a 10-gram from it; we then query each platform to find the number of search results that match that exact string. We use the custom search API for Google,7 the Bing Web Search API8 and indexes to C4 and the Pile by AI2[33].9

Table 3 lists the results of this analysis, displaying the correlation (Spearman \(\rho\)) between GPT-4 name cloze accuracy for a book and the average number of search results for all query passages from it.

Table 3: Correlation (Spearman \(\rho\)) between GPT-4 name cloze accuracy and number of search results in Google, Bing, C4 and the Pile.
Date Google Bing C4 Pile
pre-1928 0.74 0.70 0.71 0.84
post-1928 0.37 0.41 0.36 0.21

For works in the public domain (published before 1928), we see a strong and significant (\(p < 0.001\)) correlation between GPT-4 name cloze accuracy and the number of search results across all sources. Google, for instance, contains an average of 2,590 results for 10-grams from Alice in Wonderland, 1,100 results from Huckleberry Finn and 279 results from Anne of Green Gables. Works in copyright (published after 1928) show up frequently on the web as well. While the correlation is not as strong as public domain texts (in part reflecting the smaller number of copies), it is strongly significant as well (\(p < 0.001\)). Google again has an average of 3,074 search results across the ten 10-ngrams we query for Harry Potter and the Sorcerer’s Stone,10 92 results for The Hunger Games and 41 results for A Game of Thrones.

While this does not indicate direct causation between web prevalence and memorization, we speculate that there is a hidden confounder that explains them both: popularity. Moby Dick, for example, is widely read: many copies of it appear in university libraries, multiple copies of it may appear in large-scale training datasets (such as Books3, part of Pile), quotations of it may appear in unindexed academic journals, and as we see here, it appears in multiple places on the public web (both in snippets and in the form of full text). All four of these factors are likely correlated with each other and commonly explained by the overall popularity of the text itself (concretely, by an unobservable quantity such as the number of times it has been read worldwide). The correlation we see between web prevalence and memorization may be an indirect reflection of this latent causal graph.

Table 4: Sources for copyrighted material.
Domain Hits
archive.org 337
academia.edu 257
goodreads.com 234
coursehero.com 197
quizlet.com 181
litcharts.com 148
fliphtml5.com 124
genius.com 118
amazon.com 109
issuu.com 98

Table 4 lists the most popular sources for copyrighted material. Notably, this list does not only include sources where the full text is freely available for download as a single document, but also sources where smaller snippets of the texts appear in reviews (on Goodreads and Amazon) and in study resources (such as notes and quizzes on CourseHero, Quizlet and LitCharts). Since our querying process samples ten random passages from a book (not filtered according to notability or popularity in any way), this suggests that there exists a significant portion of copyrighted literary works in the form of short quotes or excerpts on the open web. On Goodreads, for example, a book can have a dedicated page of quotes from it added by users of the website (in addition to quotes highlighted in their reviews; cf. [52]); on LitCharts, the study guide of a book often includes a detailed chapter-by-chapter summary that includes quotations.

In addition, this also suggests that books that are reviewed and maintain an online presence or assigned as course readings are more likely to appear on the open web, despite copyright restrictions. In this light, observations from literary scholars and sociologies remain pertinent:  concerns, that literary taste reflects and reinforces social inequalities, could also be relevant in the context of LLMs. ’s critique on canon formation, which highlights the role of educational institutions, can still serve as a potent reminder of the impact of literary syllabi: since they have influence over what books should have flashcards and study guides created for them, they also indirectly sway which books might become more prominent on the open web, and likely in the training data for LLMs.

7 Effect on downstream tasks↩︎

Table 5: Date of first publication performance, with 95% bootstrap confidence intervals around MAE. Spearman \(\rho\) across all data and all differences between top and bottom 10% within a model are significant (\(p < 0.05\)).
missing, top 10% missing, bot. 10% \(\rho\) MAE, top 10% MAE, bot. 10%
0.04 0.09 -0.39 3.3 [0.1–12.3] 29.8 [20.2-41.2]
GPT-4 0.00 0.00 -0.37 0.6 [0.1–1.5] 14.5 [9.5–20.8]

The varying degrees to which ChatGPT and GPT-4 have seen books in their training data has the potential to bias the accuracy of downstream analyses. To test the degree to which this is true, we carry out two predictive experiments using GPT-4 in a few-shot scenario: predicting the year of first publication of a work from a 250-word passage, and predicting the amount of time that has passed within it. Note that here we are interested in the effects of disparity in memorization on downstream tasks, not any individual model’s performance on them, and for this reason, we focus on a few-shot setup without any fine-tuning.

7.1 Predicting year of first publication↩︎

Our first experiment predicts the year of first publication of a book from a random 250-word passage of text within it, a task widely used in the digital humanities when the date of composition for a work is unknown or contested[53][56].

For each of the books in our dataset, we sample a random 250-word passage (with complete sentences) and pass it through the prompt shown in figure 3 in the Appendix to solicit a four-digit date as a prediction for the year of first publication. We measure two quantities: whether a model is able make a valid year prediction at all (or refrains, noting the year is “uncertain”); if it make a valid prediction, we calculate the prediction error for each book as the absolute difference between the predicted year and the true year (\(|\hat{y} - y|\)). In order to assess any disparities in performance as a function of a model’s knowledge about a book, we focus on the differential between the books in the top 10% and bottom 10% by GPT-4 name cloze accuracy.

As table 5 shows, we see strong disparities in performance as result of a model’s familiarity with a book. ChatGPT refrains from making a prediction in only 4% of books it knows well, but 9% in books it does not (GPT-4 makes a valid prediction in all cases). When ChatGPT makes a valid prediction, the mean absolute error is much greater for the books it has not memorized (29.8 years) than for books it has (3.3 years); and the same holds for GPT-4 as well (14.5 years for non-memorized books; 0.6 years for memorized ones). Over all books with predictions, increasing GPT name cloze accuracy is linked to decreasing error rates at year prediction (Spearman \(\rho = -0.39\) for ChatGPT, \(-0.37\) for GPT-4, \(p < 0.001\)).

Predicting the date of first publication is typically cast as a textual inference problem, estimating a date from linguistic features of the text alone (e.g., mentions of airplanes may date a book to after the turn of the 20th century; thou and thee offer evidence of composition before the 18th century, etc.). While it is reasonable to believe that memorized access to an entire book may lead to improved performance when predicting the date from a passage within it, another explanation is simply that GPT models are accessing encyclopedic knowledge about those works—i.e., identifying a passage as Pride and Prejudice and then accessing a learned fact about that work. For tasks that can be expressed as facts about a text (e.g., time of publication, genre, authorship), this presents a significant risk of contamination.

7.2 Predicting narrative time↩︎

Our second experiment predicts the amount of time that has passed in the same 250-word passage sampled above, using the conceptual framework of [57], explored in the context of GPT-4 in [5]. Unlike the year prediction task above, which can potentially draw on encyclopedic knowledge that GPT-4 has memorized, this task requires reasoning directly about the information in the passage, as each passage has its own duration. Memorized information about a complete book could inform the prediction about all passages within it (e.g., through awareness of the general time scales present in a work, or the amount of dialogue within it), in addition to having exact knowledge of the passage itself.

As above, our core question asks: do we see a difference in performance between books that GPT-4 has memorized and those that it has not? To assess this, we again identify the top and bottom decile of books by their GPT-4 name cloze accuracy. We manually annotate the amount of time passing in a 250-word passage from that book using the criteria outlined by [57], including the examples provided to GPT-4 in [5], blinding annotators to the identity of the book and which experimental condition it belongs to. We then pass those same passages through GPT-4 using the prompt in figure 4, adapted from [5]. We calculate accuracy as the Spearman \(\rho\) between the predicted times and true times and calculate a 95% confidence interval over that measure using the bootstrap.

We find a large but not statistically significant difference between TOP\(_{10}\) (\(\rho = 0.50\) \([0.25\)\(0.72]\)) and BOT\(_{10}\) (\(\rho = 0.27\) \([-0.02\)\(0.59]\)), suggesting that GPT-4 may perform better on this task for books it has memorized, but not conclusive evidence that this is so. Expanding this same process to the top and bottom quintile sees no effect (TOP\(_{20}\) (\(\rho = 0.47\) \([0.29\)\(0.63]\); BOT\(_{20}\) (\(\rho = 0.46\) \([0.25\)\(0.64]\)); if an effect exists, it is only among the most memorized books. This suggests caution; more research is required to assess these disparities.

8 Discussion↩︎

We carry out this work in anticipation of the appeal of ChatGPT and GPT-4 for both posing and answering questions in cultural analytics. We find that these models have memorized books, both in the public domain and in copyright, and the capacity for memorization is tied to a book’s overall popularity on the web. This differential in memorization leads to differential in performance for downstream tasks, with better performance on popular books than on those not seen on the web. As we consider the use of these models for questions in cultural analytics, we see a number of issues worth considering.

8.0.0.1 The virtues of open models.

This archaeology is required only because ChatGPT and GPT-4 are closed systems, and the underlying data on which they have been trained is unknown. While our work sheds light on the memorization behavior of these models, the true training data remains fundamentally unknowable outside of OpenAI. As others have pointed out with respect to these systems in particular, there are deep ethical and reproducibility concerns about the use of such closed models for scholarly research[58]. A preferred alternative is to embrace the development of more open, transparent models, including LLaMA[59], OPT[60] and BLOOM[61]. While open-data models will not alleviate the disparities in performance we see by virtue of their openness (LLaMA is trained on books in the Pile, and all use information from the general web, which each allocate attention to popular works), they address the issue of data contamination since the works in the training data will be known; works in any evaluation data would then be able to be queried for membership[62].

8.0.0.2 Popular texts are likely not good barometers of model performance.

A core concern about closed models with unknown training data is test contamination: data in an evaluation benchmark (providing an assessment of measurement validity) may be present in the training data, leading an assessment to be overconfident in a model’s abilities. Our work here has shown that OpenAI models know about books in proportion to their popularity on the web, and that their performance on downstream tasks is tied to that popularity. When benchmarking the performance of these models on a new downstream task, it is risky to draw expectations of generalization from its performance on Alice in Wonderland, Harry Potter, Pride and Prejudice, and so on—it simply knows much more about these works than the long tail of literature (both in the public domain and in copyright). In this light, we hope this work can serve as the starting point for future work on the generalization ability of LLM in the context of cultural analytics.

8.0.0.3 Whose narratives?

At the core of questions surrounding the data used to train large language models is one of representation: whose language, and whose lived experiences mediated through that language, is captured in their knowledge of the world? While previous work has investigated this in terms of what varieties of English pass GPT-inspired content filters[63], we can ask the same question about the narrative experiences present in books: whose narratives inform the implicit knowledge of these models? Our work here has not only confirmed that ChatGPT and GPT-4 have memorized popular works, but also what kinds of narratives count as “popular”—in particular, works in the public domain published before 1928, along with contemporary science fiction and fantasy. Our work here does not investigate how this potential source of bias might become realized—e.g., how generated narratives are pushed to resemble text it has been trained on[64]; how the values of the memorized texts are encoded to influence other behaviors, etc.—and so any discussion about it can only be speculative, but future work in this space can take the disparities we find in memorization as a starting point for this important work.

9 Conclusion↩︎

As research in cultural analytics is poised to be transformed by the affordances of large language models, we carry out this work to uncover the works of fiction that two popular models have memorized, in order to assess the risks of using closed models for research. In our finding that memorization affects performance on downstream tasks in disparate ways, closed data raises uncertainty about the conditions when a model will perform well and when it will fail; while our work tying memorization to popularity on the web offers a rough guide to assess the likelihood of memorization, it does not solve that fundamental challenge. Only open models with known training sources would do so.

Code and data to support this work can be found at https://github.com/bamman-group/gpt4-books.

Limitations↩︎

The data behind ChatGPT and GPT-4 is fundamentally unknowable outside of OpenAI. Our work carries out probabilistic inference to measure the familiarity of these models with a set of books, but the question of whether they truly exist within the training data of these models is not answerable. Mere familiarity, however, is enough to contaminate test data.

Our study has some limitations that future work could address. First, we focus on a case study of English, but the analytical method we employ is transferable to other languages (given digitized books written in those languages). We leave it to future work to quantify the extent of memorization in non-English literary content by large language models.

Second, the name cloze task is one operationalization to assess memorization of large language models. As shown by previous research [41], memorization can be much more acute with longer textual spans during training copied verbatim. We leave quantifying the full extent of memorization and data contamination of opaque models such as ChatGPT and GPT-4 for future work.

Ethical considerations↩︎

Our study probes the degree to which ChatGPT and GPT-4 have memorized books, and what impact that memorization has on downstream research in cultural analytics. This work uses the OpenAI API to perform the name cloze task described above, and at no point do we access, or attempt to access, the true training data behind these models, or any underlying components of the systems.

Acknowledgments↩︎

The research reported in this article was supported by funding from the National Science Foundation (IIS-1942591) and the National Endowment for the Humanities (HAA-271654-20).

Table 6: Top 50 books in copyright (published after 1928) by GPT-4 name cloze accuracy.
GPT-4 ChatGPT BERT Date Author Title
0.76 0.43 0.00 1997 J.K. Rowling Harry Potter and the Sorcerer’s Stone
0.57 0.30 0.00 1949 George Orwell 1984
0.51 0.20 0.01 1954 J.R.R. Tolkien The Fellowship of the Ring
0.49 0.16 0.13 2012 E.L. James Fifty Shades of Grey
0.48 0.14 0.00 2008 Suzanne Collins The Hunger Games
0.43 0.27 0.00 1954 William Golding Lord of the Flies
0.43 0.17 0.00 1979 Douglas Adams The Hitchhiker’s Guide to the Galaxy
0.30 0.16 0.00 1959 Chinua Achebe Things Fall Apart
0.28 0.12 0.00 1977 J. R. R. Tolkien and Christopher Tolkien The Silmarillion
0.27 0.13 0.00 1953 Ray Bradbury Fahrenheit 451
0.27 0.13 0.00 1996 George R.R. Martin A Game of Thrones
0.26 0.05 0.01 2003 Dan Brown The Da Vinci Code
0.26 0.08 0.00 1965 Frank Herbert Dune
0.25 0.20 0.01 1937 Zora Neale Hurston Their Eyes Were Watching God
0.25 0.14 0.00 1961 Harper Lee To Kill a Mockingbird
0.24 0.03 0.03 1953 Ian Fleming Casino Royale
0.22 0.13 0.00 1984 William Gibson Neuromancer
0.20 0.10 0.00 1985 Orson Scott Card Ender’s Game
0.19 0.12 0.00 1932 Aldous Huxley Brave New World
0.18 0.07 0.00 1937 Margaret Mitchell Gone with the Wind
0.17 0.05 0.00 1968 Philip K. Dick Do Androids Dream of Electric Sheep?
0.16 0.06 0.05 2009 Dan Brown The Lost Symbol
0.15 0.04 0.04 2013 Dan Brown Inferno
0.15 0.08 0.01 2014 Veronica Roth Divergent
0.15 0.04 0.00 1940 John Steinbeck The Grapes of Wrath
0.13 0.05 0.00 1983 James Kahn Return of the Jedi
0.13 0.02 0.01 1928 D. H. Lawrence Lady Chatterley’s Lover
0.13 0.03 0.00 1977 Alex Haley Roots
0.11 0.11 0.00 1961 Irving Stone The Agony and the Ecstasy
0.11 0.01 0.03 1957 Ian Fleming From Russia with Love
0.11 0.04 0.00 1962 Madeleine L’Engle A Wrinkle in Time
0.11 0.04 0.01 1939 Marjorie Kinnan Rawlings The Yearling
0.10 0.05 0.00 1975 E. L. Doctorow Ragtime
0.10 0.05 0.00 1929 Dashiell Hammett The Maltese Falcon
0.10 0.08 0.07 1991 Diana Gabaldon Outlander
0.10 0.02 0.00 1989 Kazuo Ishiguro The Remains of the Day
0.10 0.01 0.00 1983 Alice Walker The Color Purple
0.09 0.02 0.00 1934 Dorothy L. Sayers The Nine Tailors
0.09 0.03 0.00 1985 Margaret Atwood The Handmaid’s Tale
0.09 0.04 0.00 1988 Toni Morrison Beloved
0.08 0.07 0.00 1982 Joe Nazel Every Goodbye Ain’t Gone
0.08 0.06 0.01 1984 Helen Hooven Santmyer ...And Ladies of the Club
0.08 0.02 0.00 2006 Max Brooks World War Z
0.08 0.02 0.00 1993 Irvine Welsh Trainspotting
0.08 0.03 0.00 1947 Robert Penn Warren All the King’s Men
0.07 0.02 0.00 1952 Ralph Ellison Invisible Man
0.07 0.01 0.00 1951 Isaac Asimov Foundation
0.07 0.01 0.00 1976 Anne Rice Interview with the Vampire
0.07 0.04 0.02 1977 Stephen King The Shining
0.07 0.01 0.00 1996 Helen Fielding Bridget Jones’s Diary

Appendix↩︎

Author contributions↩︎

  • Kent Chang: performed extrinsic analysis; annotated narrative time duration data; wrote paper.

  • Mackenzie Cramer: performed human name cloze experiment; annotated narrative time duration data; researched self-publishing history of Fifty Shades of Grey.

  • Sandeep Soni: performed the statistical analysis of predicted names; annotated narrative time duration data; wrote paper.

  • David Bamman: conducted experiments, annotated narrative time duration data; wrote paper.

None

Figure 3: Sample prompt for year of first publication prediction..

None

Figure 4: Sample prompt for narrative time prediction, from [5]..

References↩︎

[1]
Andrew Piper, Richard Jean So, and David Bamman. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.26. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 298–311, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
[2]
Michael Yoder, Sopan Khosla, Qinlan Shen, Aakanksha Naik, Huiming Jin, Hariharan Muralidharan, and Carolyn Rosé. 2021. https://doi.org/10.18653/v1/2021.nuse-1.2. In Proceedings of the Third Workshop on Narrative Understanding, pages 13–23, Virtual. Association for Computational Linguistics.
[3]
Mariona Coll Ardanuy, Federico Nanni, Kaspar Beelen, Kasra Hosseini, Ruth Ahnert, Jon Lawrence, Katherine McDonough, Giorgia Tolfo, Daniel CS Wilson, and Barbara McGillivray. 2020. https://doi.org/10.18653/v1/2020.coling-main.400. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4534–4545, Barcelona, Spain (Online). International Committee on Computational Linguistics.
[4]
Elizabeth F. Evans and Matthew Wilkens. 2018. Nation, ethnicity, and the geography of British fiction, 1880–1940. Cultural Analytics.
[5]
Ted Underwood. 2023. Using GPT-4 to measure the passage of time in fiction. https://tedunderwood.com/2023/03/19/using-gpt-4-to-measure-the-passage-of-time-in-fiction/.
[6]
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. https://aclanthology.org/2022.findings-emnlp.59. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 840–854, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
[7]
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022. Large language models struggle to learn long-tail knowledge. arXiv preprint arXiv:2211.08411.
[8]
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Amir Feder, Abhilasha Ravichander, Marius Mosbach, Yonatan Belinkov, Hinrich Schütze, and Yoav Goldberg. 2022. https://arxiv.org/abs/2207.14251. arXiv preprint arXiv:2207.14251.
[9]
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022. https://arxiv.org/abs/2202.07646. arXiv preprint arXiv:2202.07646.
[10]
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023. https://arxiv.org/abs/2304.01373. arXiv preprint arXiv:2304.01373.
[11]
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets. Communications of the ACM, 64(12):86–92.
[12]
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP), pages 3–18. IEEE.
[13]
Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022. Memorization without overfitting: Analyzing the training dynamics of large language models. Advances in Neural Information Processing Systems, 35:38274–38290.
[14]
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang. 2023. Foundation models and fair use. arXiv preprint arXiv:2303.15715.
[15]
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023. https://arxiv.org/abs/2301.13188. arXiv preprint arXiv:2301.13188.
[16]
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pages 19–27.
[17]
David M. Berry. 2011. The computational turn: Thinking about the digital humanities. Culture Machine, 12.
[18]
Kathleen Fitzpatrick. 2012. https://books.google.com/books?hl=en&lr=&id=_6mo2tApzQQC&oi=fnd&pg=PA12&dq=The+Humanities+Done+Digitally&ots=9ZBdlpM8Ty&sig=1erKMiLDk3zQVLn2ArUhNQHTBuQ. Debates in the digital humanities, pages 12–15.
[19]
N. Katherine Hayles. 2012. https://doi.org/10.1057/9780230371934_3. In David M Berry, editor, Understanding Digital Humanities, pages 42–66. Palgrave Macmillan UK, London.
[20]
Stephen Ramsay and Geoffrey Rockwell. 2012. https://books.google.com/books?hl=en&lr=&id=_6mo2tApzQQC&oi=fnd&pg=PA75&dq=Developing+Things+Notes+toward+an+Epistemology+of+Building+in+the+Digital+Humanities&ots=9ZBdlpN6Mx&sig=0q0-_3Wvo5a-MC0jSwojh7MW-gk. Debates in the digital humanities, pages 75–84.
[21]
Nicolas Ruth, Andreas Niekler, and Manuel Burghardt. 2022. https://ceur-ws.org/Vol-3290/long_paper6029.pdf. Proceedings of the Computational Humanities Research Conference 2022, 1613:0073.
[22]
Katherine Elkins and Jon Chun. 2020. https://culturalanalytics.org/article/17212.pdfJournal of cultural analytics, 5(2).
[23]
Lauren M E Goodlad and Wai Chee Dimock. 2021. https://www.cambridge.org/core/journals/pmla/article/ai-and-the-human/3FEBDB0945D8CF022EC949D945C80CEA. PMLA, 136(2):317–319.
[24]
Leah Henrickson and Albert Meroño-Peñuela. 2022. https://muse.jhu.edu/article/853606. Configurations, 30(2):115–139.
[25]
Henning Schmidgen, Bernhard Dotzler, and Benno Stein. 2023. https://culturalanalytics.org/article/55795-from-the-archive-to-the-computer-michel-foucault-and-the-digital-humanities. Journal of cultural analytics, 7(4).
[26]
M Elam. 2023. https://read.dukeupress.edu/american-literature/article-abstract/doi/10.1215/00029831-10575077/344231?casa_token=16aoCnivZ74AAAAA:R8hlKO7e2pPTNg3WbaFMyexxOEzkl9s-hz9VPAsKqZooRHCoMTjutjRJVC2mQuoq5-Q2-kVXAmerican literature; a journal of literary history, criticism and bibliography.
[27]
Gary Gutting. 1989. https://play.google.com/store/books/details?id=JuOkN9gSP04C. Cambridge University Press.
[28]
Sil Hamilton and Andrew Piper. 2022. https://aclanthology.org/2022.latechclfl-1.11. In Proceedings of the 6th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, pages 83–93, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
[29]
Dominik Stammbach, Maria Antoniak, and Elliott Ash. 2022. https://doi.org/10.18653/v1/2022.wnu-1.6. In Proceedings of the 4th Workshop of Narrative Understanding (WNU2022), pages 47–56, Seattle, United States. Association for Computational Linguistics.
[30]
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2023. Can large language models transform computational social science? arXiv submission 4840038.
[31]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
[32]
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. https://arxiv.org/abs/2101.00027. arXiv preprint arXiv:2101.00027.
[33]
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.98. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1286–1305, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
[34]
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.746. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9275–9293, Online. Association for Computational Linguistics.
[35]
Aparna Elangovan, Jiayuan He, and Karin Verspoor. 2021. https://doi.org/10.18653/v1/2021.eacl-main.113. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1325–1335, Online. Association for Computational Linguistics.
[36]
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021. https://doi.org/10.18653/v1/2021.eacl-main.86. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1000–1008, Online. Association for Computational Linguistics.
[37]
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
[38]
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. https://www.usenix.org/system/files/sec21-carlini-extracting.pdf In USENIX Security Symposium, volume 6.
[39]
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022. Are large pre-trained language models leaking your personal information? arXiv preprint arXiv:2205.12628.
[40]
Fatemehsadat Mireshghallah, Archit Uniyal, Tianhao Wang, David Evans, and Taylor Berg-Kirkpatrick. 2022. https://aclanthology.org/2022.emnlp-main.119. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1816–1826, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
[41]
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022. https://doi.org/10.18653/v1/2022.acl-long.577. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, Dublin, Ireland. Association for Computational Linguistics.
[42]
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A Choquette-Choo, and Nicholas Carlini. 2022. Preventing verbatim memorization in language models gives a false sense of privacy. arXiv preprint arXiv:2210.17546.
[43]
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2021. Counterfactual memorization in neural language models. arXiv preprint arXiv:2112.12938.
[44]
Inbal Magar and Roy Schwartz. 2022. https://doi.org/10.18653/v1/2022.acl-short.18. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 157–165, Dublin, Ireland. Association for Computational Linguistics.
[45]
Terra Blevins and Luke Zettlemoyer. 2022. https://aclanthology.org/2022.emnlp-main.233. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3563–3574, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
[46]
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022. Deduplicating training data mitigates privacy risks in language models. In International Conference on Machine Learning, pages 10697–10707. PMLR.
[47]
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015. The Goldilocks principle: Reading children’s books with explicit memory representations. ICLR.
[48]
Takeshi Onishi, Hai Wang, Mohit Bansal, Kevin Gimpel, and David McAllester. 2016. https://doi.org/10.18653/v1/D16-1241. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2230–2235, Austin, Texas. Association for Computational Linguistics.
[49]
Hila Gonen, Srini Iyer, Terra Blevins, Noah A. Smith, and Luke Zettlemoyer. 2022. http://arxiv.org/abs/2212.04037.
[50]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
[51]
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018. https://doi.org/10.18653/v1/P18-1128. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1383–1392, Melbourne, Australia. Association for Computational Linguistics.
[52]
Melanie Walsh and Maria Antoniak. 2021. The Goodreads Classics”: A computational study of readers, Amazon, and crowdsourced amateur criticism. Journal of Cultural Analytics, 6(2).
[53]
FM Jong, Henning Rode, and Djoerd Hiemstra. 2005. Temporal language models for the disclosure of historical text. Proceedings of the XVIth International Conference of the Association for History and Computing (AHC 2005).
[54]
Abhimanu Kumar, Matthew Lease, and Jason Baldridge. 2011. Supervised language modeling for temporal resolution of texts. In Proceedings of the 20th ACM international conference on Information and knowledge management, pages 2069–2072.
[55]
David Bamman, Michelle Carney, Jon Gillick, Cody Hennesy, and Vijitha Sridhar. 2017. https://ieeexplore.ieee.org/document/7991569. In 2017 ACM/IEEE Joint Conference on Digital Libraries (JCDL), pages 1–10. IEEE.
[56]
Folgert Karsdorp and Lauren Fonteyn. 2019. Cultural entrenchment of folktales is encoded in language. Palgrave Communications, 5(1):25.
[57]
Ted Underwood. 2018. Why literary time is measured in minutes. ELH, 85(2).
[58]
Arthur Spirling. 2023. Why open-source generative AI models are an ethical way forward for science. Nature, 616(7957):413–413.
[59]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. https://arxiv.org/abs/2302.13971. arXiv preprint arXiv:2302.13971.
[60]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
[61]
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. https://arxiv.org/abs/2211.05100. arXiv preprint arXiv:2211.05100.
[62]
Marc Marone and Benjamin Van Durme. 2023. https://doi.org/10.48550/ARXIV.2303.03919.
[63]
Suchin Gururangan, Dallas Card, Sarah K Drier, Emily K Gade, Leroy Z Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A Smith. 2022. Whose language counts as high quality? Measuring language ideologies in text data selection. arXiv preprint arXiv:2201.10474.
[64]
Li Lucy and David Bamman. 2021. https://doi.org/10.18653/v1/2021.nuse-1.5. In Proceedings of the Third Workshop on Narrative Understanding, pages 48–55, Virtual. Association for Computational Linguistics.

  1.   Details of author contributions listed in the appendix.↩︎

  2. https://github.com/booknlp/booknlp↩︎

  3. http://bbip.ku.edu/novel-collections↩︎

  4. For complete results on all books, see https://github.com/bamman-group/gpt4-books.↩︎

  5. Fifty Shades of Grey was originally self-published on fanfiction.net ca. 2009 before being published by Vintage Books in 2012.↩︎

  6. https://huggingface.co/datasets/bookcorpus↩︎

  7. https://developers.google.com/custom-search/v1/overview↩︎

  8. https://www.microsoft.com/en-us/bing/apis/bing-web-search-api↩︎

  9. https://c4-search.apps.allenai.org↩︎

  10. cf. “seemed to know where he was going, he was obviously”; “He bent and pulled the ring of the trapdoor, which”; “I got a few extra books for background reading, and”↩︎