<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">2940-1348</journal-id>
<journal-title-group>
<journal-title>Journal of Computational Literary Studies</journal-title>
</journal-title-group>
<issn pub-type="epub">2940-1348</issn>
<publisher>
<publisher-name>Technische Universit&#228;t Darmstadt</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.48694/jcls.3927</article-id>
<article-categories>
<subj-group>
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>From Review to Genre to Novel and Back</article-title>
<subtitle>An Attempt to Relate Reader Impact to Phenomena of Novel Text</subtitle>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-0301-2029</contrib-id>
<name>
<surname>Koolen</surname>
<given-names>Marijn</given-names>
</name>
<email>marijn.koolen@gmail.com</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-3862-7602</contrib-id>
<name>
<surname>van Zundert</surname>
<given-names>Joris J.</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-1330-0585</contrib-id>
<name>
<surname>Viviani</surname>
<given-names>Eva</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-9139-1577</contrib-id>
<name>
<surname>Schnober</surname>
<given-names>Carsten</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-6478-3003</contrib-id>
<name>
<surname>van Hage</surname>
<given-names>Willem</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-1916-1851</contrib-id>
<name>
<surname>Tereshko</surname>
<given-names>Katja</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>DHLab, Humanities Cluster, Amsterdam, The Netherlands</aff>
<aff id="aff-2"><label>2</label>Computational Literary Research, Huygens Institute <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://ror.org/04x6kq749">ROR</ext-link>, Amsterdam, The Netherlands</aff>
<aff id="aff-3"><label>3</label>eScience Center <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://ror.org/00rbjv475">ROR</ext-link>, Amsterdam, The Netherlands</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2024-10-17">
<day>17</day>
<month>10</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>3</volume>
<issue>1</issue>
<fpage>1</fpage>
<lpage>31</lpage>
<history>
<date date-type="received" iso-8601-date="2024-01-25">
<day>25</day>
<month>01</month>
<year>2024</year>
</date>
<date date-type="accepted" iso-8601-date="2024-09-28">
<day>28</day>
<month>09</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2024 The Author(s)</copyright-statement>
<copyright-year>2024</copyright-year>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>The text of this work is released under the Creative Commons license CC BY 4.0 International. You can find the contract text of the license at <uri xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</uri>. The illustrations are excluded from this license, here the copyright lies with the respective rights holder.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://jcls.io/articles/10.48694/jcls.3927/"/>
<abstract>
<p>We are interested in the textual features that correlate with the reported impact by readers of novels. We operationalize impact measurement through a rule-based reading impact model and apply it to 634,614 reader reviews mined from seven review platforms. We compute co-occurrences of impact-related terms and their keyness for genres represented in the corpus. The corpus consists of the full text of 18,885 books from which we derived topic models. The topics we find correlate strongly with genre, and we get strong indicators for which key impact terms are connected to which genre. These key impact terms give us a first evidence-based insight into genre-related readers&#8217; motivations.</p>
</abstract>
<kwd-group>
<kwd>reading impact</kwd>
<kwd>literary novels</kwd>
<kwd>genre</kwd>
<kwd>topic modeling</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="S1">
<title>1. Introduction</title>
<p>Already Aristotle noted the reciprocal relations between an author, the text the author creates, and the response from an audience to the text. This fundamental model of rhetorical poetics has remained relevant throughout the ages (see e.g., <xref ref-type="bibr" rid="B1">Abrams 1971</xref>; <xref ref-type="bibr" rid="B41">Warnock 1978</xref>). The dynamics of the relations between author, text, and reader have been heavily theorized and fiercely debated (see e.g., <xref ref-type="bibr" rid="B21">Hickman 2012</xref>; <xref ref-type="bibr" rid="B42">Wimsatt 1954</xref>). But if there is no lack of theory, it appears to be much harder to gain empirical insights into these relations, though not for lack of trying by practitioners in such fields as empirical and computational literary studies (e.g., <xref ref-type="bibr" rid="B17">Fialho 2019</xref>; <xref ref-type="bibr" rid="B26">Loi et al. 2023</xref>; <xref ref-type="bibr" rid="B28">Miall and Kuiken 1994</xref>). One effect of the immense success of the World Wide Web and softwarization and digitization of societies and their cultures (<xref ref-type="bibr" rid="B5">Berry 2014</xref>; <xref ref-type="bibr" rid="B27">Manovich 2013</xref>) is the availability of large collections of online book reviews and digital full texts from novels published as ePubs. This allows us to apply NLP techniques and corpus statistics to get empirical data on the relations between text and reader that until now could only be theorized or anecdotally evidenced. At the same time, we should acknowledge that it is no panacea for the problem of empirical observations in literary studies. Not just because of the inherent biases (<xref ref-type="bibr" rid="B19">Gitelman 2013</xref>; <xref ref-type="bibr" rid="B33">Prescott 2023</xref>; <xref ref-type="bibr" rid="B34">Rawson and Mu&#241;oz 2016</xref>), or the almost complete lack of demographic and social signals in the data, but also because of the difficulties still involved in establishing which concrete signal in novels relates to which type of reaction for which type of reader. This is where we focus our research: We attempt to establish which concrete features of online reviews correlate to which concrete signals in the text of fiction novels.</p>
<p>&#921;n a theoretical sense, we are concentrating on the right hand side of the classical rhetorical triangle (see <xref ref-type="fig" rid="F1">Figure 1a</xref>) and operationalize the dynamic between text and reader as another triangular relationship between <italic>impact, topic</italic>, and <italic>genre</italic>. With &#8220;impact&#8221; (and the commensurate &#8220;reading impact&#8221;), we designate expressions of reader experiences identified by some evidence-based method (e.g., as reader impact constituents researched by Koolen et al. (<xref ref-type="bibr" rid="B23">2023</xref>)). We apply the reader impact model to assign concrete terms to types of reading impact. The concrete text signal that we correlate this impact with are topics mined from a corpus of novels. (As an aside, we note that these topics are not to be confused with themes, motives, or aboutness in a literary studies sense, as we will explain later.) The meta-textual property, genre, forms the third measurable aspect of the triangular relationship (see <xref ref-type="fig" rid="F1">Figure 1b</xref>).</p>
<fig id="F1">
<caption>
<p><bold>Figure 1:</bold> Classic rhetorical model (a) and our operationalization of the text&#8211;reader relation (b).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g1.png"/>
</fig>
<p>Concretely, we link topic models of 18,885 novels in Dutch (original Dutch and translated to Dutch) with the reading impact expressed in 130,751 Dutch online book reviews. We want to know if there is a relationship between aspects of topic in novels, their genre, and the type of impact expressed by readers in their reviews. We extracted expressions for three types of reading impact from the reviews using the previously developed Reading Impact Model for Dutch (<xref ref-type="bibr" rid="B9">Boot and Koolen 2020</xref>). The three types of reading impact that we discern are: &#8220;General affective impact&#8221;, which expresses the overall evaluation and sentiment regarding a novel; &#8220;narrative impact&#8221;, which relates to aspects of story, plot, and characters; and finally &#8220;stylistic impact&#8221;, related to writing style and aesthetics.</p>
<p>We expect that topics in fiction are related to genre. As there is no authoritative source for genre of a novel, nor some general academic consensus about what constitutes genre, we make use of the broad genre labels that publishers have assigned to each published book. Analogous to Sobchuk and &#352;e&#316;a (<xref ref-type="bibr" rid="B37">2023, 2</xref>), who define genre as &#8220;a population of texts united by broad thematic similarities&#8221;, we clustered these genre labels into a set of nine genres. These thematic similarities might be revealed in a topical analysis, e.g., crime novels containing more crime-related topics and romance novels containing more topics related to romance and sex. However, for some genres it might be less obvious whether they are related to topic. For instance, what are the topics one would expect in the broad genre of literary fiction?</p>
<p>It is important to note that, although the name <italic>topic modeling</italic> suggests that what is modeled is <italic>topic</italic>, most topic modeling approaches discern clusters of frequently co-occurring words, regardless of whether they have a topical connection or not (in the classical sense of &#8220;aboutness&#8221; in library science). Clusters of words may also reveal a different type of connection, e.g., words from a particular stylistic register. In that sense, genres with less clear thematic similarities may be associated with certain stylistic registers, or any other clustering of vocabulary. Different genres may also attract different types of readers and therefore different types of reviewers, who use different terminology and pay attention to different aspects of novels. It is also plausible that the language and topic of a novel influences how readers write about them in reviews. A novel written in a particularly striking poetic style may consciously or subconsciously lead readers to adopt some of its poetic aspects and register in how they write about their reading experiences. Similarly, topics in novels may be associated with what reviewers choose to mention, again, consciously or subconsciously. A novel on the atrocities of war or on the pain of losing a loved one may lead a reviewer to mention feeling sympathy or sadness during reading, while a story about friendship and betrayal might prompt reviewers to describe their anger at the actions of one of the characters.</p>
<p>Thus, it is clear that the relationship between the three elements &#8211; topic, genre, and impact &#8211; is complex and reciprocal, as expressed in <xref ref-type="fig" rid="F1">Figure 1b</xref>. Our challenge, of course, is to computationally investigate and understand this relationship utilizing the large number of full-text novels from different genres and corpora of hundreds of thousands of reviews. We subdivide this overarching aim into several more concrete research questions, namely:</p>
<list list-type="bullet">
<list-item><p>How are topic and impact related to each other? Do books with certain topics lead to more impact expressed in book reviews? Do different topics lead to different types of impact?</p></list-item>
<list-item><p>How are genre and impact related to each other? Do books of different genres lead to different types of impact? Do reviews of different genres use different vocabulary for expressing the same types of impact?</p></list-item>
<list-item><p>How are topic and genre related to each other? Are certain topics more likely in some genres than in others?</p></list-item>
</list>
<p>This paper makes three main contributions to our ongoing research. The first is that it contributes to our understanding of the reading impact model and through it of the language of reading impact. We formalize the ability to tell genres apart using the <italic>keyness</italic> of impact terms. Thus, we now have quantitative support to argue that certain impact terms are strongly connected to certain genres and less to others. Second, we find that the topics from novels can be clustered into broader themes that lead to distinct thematic profiles per genre. There is a clear relation between impact terms and genre, but not between impact terms and topic or theme. In the discussion at the end, we elaborate on this and provide possible explanations for this finding. The third contribution is the insight that the key impact terms per genre give an indication of the motivation of readers to read a book and how the reading experience relates to their expectations.</p>
</sec>
<sec id="S2">
<title>2. Background</title>
<p>We are interested in what kind of impression novels leave with their readers. Can we measure this so-called &#8220;impact&#8221; and how does it relate to features of the actual novel texts? Several studies have tried to link success or popularity of texts to features of those texts. Some studies have related pace, in the sense of how much distance the same length of texts covers in a semantic space, to success; finding that success correlates with higher pacing of narrative (<xref ref-type="bibr" rid="B24">Laurino Dos Santos and Berger 2022</xref>; <xref ref-type="bibr" rid="B39">Toubia et al. 2021</xref>). It has been argued that songs whose lyrics deviate form a genre&#8217;s usual pattern tend to be more popular (<xref ref-type="bibr" rid="B4">Berger and Packard 2018</xref>). Other work relating topic models to surveyed ratings of literariness suggests the same for fiction novels (<xref ref-type="bibr" rid="B11">Cranenburgh et al. 2019</xref>). Moreira et al. (<xref ref-type="bibr" rid="B29">2023, 32</xref>) apply &#8220;sentiment arc features [&#8230;] and semantic profiling&#8221; with some success to predict ratings on Goodreads. Taking the number of Gutenberg downloads as a proxy for success, Ashok et al. (<xref ref-type="bibr" rid="B3">2013</xref>) reach 84% accuracy in predicting popularity based on learning low-level stylistic features of the text of novels. Zundert et al. (<xref ref-type="bibr" rid="B43">2018</xref>) use sales numbers as a proxy for popularity in a machine learning attempt to predict success, concluding that the theme of masculinity is at least one major driver of successful fiction.</p>
<p>Common to all these studies is that they target some proxy of success or popularity: Goodreads ratings, sales numbers, download statistics, and so forth. However, to our knowledge no research has tried to link concrete features of fiction narratives to textual features of reviews from readers. We seek to uncover if there is such a relation and if it may be meaningful from a literary research perspective. In our present study, we apply a heuristic model for impact features (<xref ref-type="bibr" rid="B9">Boot and Koolen 2020</xref>) to a corpus of 600,000+ reader reviews mined from several online review platforms. We attempt to relate collocations of impact related terms to genre. Advancing previous research on genre and topic models (<xref ref-type="bibr" rid="B44">Zundert et al. 2022</xref>), our contribution in this paper is to examine how collocated impact terms relate to genre and genre to topic models of novels, thus offering a first insight into the relation between topics (understood in terms of topic model) and reader reported impact measures. Such work needs to take into account the plethora of problems that surround the application of topic models to downstream tasks. This concerns topics content wise, which is to say that topic models in contrast to their name do not often express much topical information. Rather they may be connected to meta-textual features, such as author (<xref ref-type="bibr" rid="B38">Thompson and Mimno 2018</xref>), genre (<xref ref-type="bibr" rid="B36">Sch&#246;ch 2017</xref>), or structural elements in texts (<xref ref-type="bibr" rid="B40">Uglanova and Gius 2020</xref>).</p>
<p>Our current contribution leans more to the side of data exploration than to the side of offering assertive generalizations. We are interested in empirically quantifying the impact that the text of novels has on readers. Any operationalization of this research aim necessarily involves many narrowing choices and, at least initially, the audacious naivety to ignore the stupefying complexity of social mechanisms to which readers are susceptible and thus the mass of confounding text-external factors that also drive reader impact. In our setup, we assume that there are at least some textual features, such as style, narrative pace, plot, character likability, that may be measured and that can be related to reader impact. We further assume that book reviews scraped from online platforms do serve as a somewhat reliable gauge to measure reader impact. We make these cautionary statements not just <italic>pro forma</italic>, but because we know that our information is selective, biased, and skewed. Thanks to the stalwart experts of the Dutch National Library, we do have for our analysis the full text of 18,885 novels in Dutch (both translated and of Dutch origin). We also have 634,614 online reviews, gathered by scraping platforms such as Goodreads, Hebban<xref ref-type="fn" rid="n1">1</xref>, and so forth. This corpus is biased. Romance novels comprise only about 3% of the corpus of full texts. This is in stark contrast to its undisputed popularity (see Regis (<xref ref-type="bibr" rid="B35">2003, xi</xref>): &#8220;In the last year of the twentieth century, 55.9% of mass-market and trade paperbacks sold in North America were romance novels&#8221;). If our book corpus is skewed, our review data is even more so: Only 1% of the reviews pertain to novels in the romance genre. Obviously, we attempt to balance our data with respect to genre and other properties for analysis. Yet we should remind ourselves of the limited representativeness of our data, which necessitates modesty as to generalizing results. Hence, what follows is more offered as data exploration than as pontification of strong relations.</p>
</sec>
<sec id="S3">
<title>3. Data and Method</title>
<p>Our corpus of 18,885 books consists of mostly fiction novels and some <italic>non-fiction</italic> books in the Dutch language (both originally Dutch and translated). The review corpus boasts 634,614 Dutch book reviews. Obviously, we do not have reviews for each book, nor does the set of books fully cover the collection of reviews, but we have upward of 10,000 books with at least one review.</p>
<sec id="S3.1">
<title>3.1 Preprocessing</title>
<p>Both &#8211; books and reviews &#8211; are parsed with Trankit (<xref ref-type="bibr" rid="B30">Nguyen et al. 2021</xref>). Reading impact is extracted from the reviews using the Dutch Reading Impact Model (DRIM) (<xref ref-type="bibr" rid="B9">Boot and Koolen 2020</xref>).</p>
<p><bold>Topic modeling</bold> For topic modeling of the novels, we use Top2Vec (<xref ref-type="bibr" rid="B2">Angelov 2020</xref>), and created a model with whole books as documents. We apply multiple filters to select terms that signal a topic. Following the advice from previous work (<xref ref-type="bibr" rid="B37">Sobchuk and &#352;e&#316;a 2023</xref>; <xref ref-type="bibr" rid="B40">Uglanova and Gius 2020</xref>; <xref ref-type="bibr" rid="B44">Zundert et al. 2022</xref>), we focus on content words, and select only nouns, verbs, adjectives, and adverbs, and remove any person names identified by the Trankit NER tagger. Our assumption is that person names have little to no relationship with topic, but are strong differentiating terms that tend to cluster parts of books and book series with recurring characters. Names of locations can have a similar effect, but, at least where the setting reflects the real world, we argue that this setting aspect of stories is more meaningfully related to topic. The book corpus contains 1,922,833,614 tokens, including all punctuation and stop words. After filtering for person and location names, 826,226,855 tokens remain.</p>
<p>The next filter is a frequency filter. We remove terms that occur in fewer than 1% of documents or in more than 50% of documents. This leaves 190,607,470 tokens which is 23% of all content words and just under 10% of the total number of tokens<xref ref-type="fn" rid="n2">2</xref>. Books have a mean (median) number of 42,959 (37,940) <italic>content</italic> tokens. The number of tokens is a Poisson distribution, therefore left-skewed, with 68% (corresponding to data within 1 standard deviation from the mean) of all books having between 17,509 and 63,418 tokens. This shows that the books have a high variation in length, but the majority of the books have a length within a single order of magnitude. After filtering on document frequency, the mean (median) number of tokens is 9,979 (8,325), with 68% having between 3,847 and 14,992 tokens.</p>
<p><bold>Reading impact modeling</bold> The DRIM is a rule-based model and works at the level of sentences. It has 275 rules relating to impact in four categories: <italic>Affect, Aesthetic</italic> and <italic>Narrative</italic> impact, and <italic>Reflection</italic>. Both <italic>Aesthetic</italic> and <italic>Narrative</italic> impact are sub-categories of <italic>Affect</italic>, so rules that identify expressions of the sub-categories are also considered expressions of <italic>Affect</italic> (<xref ref-type="bibr" rid="B9">Boot and Koolen 2020</xref>), but expressions of <italic>Affect</italic> are not necessarily counted as one of the subcategories. The rules for <italic>Reflection</italic> were not validated (see <xref ref-type="bibr" rid="B9">Boot and Koolen 2020</xref>), so we exclude <italic>Reflection</italic> from our analysis. For our analysis of topic, we expect that <italic>Narrative</italic> is the most directly related category, but we also include general <italic>Affect</italic> in our analysis. Expressions identified by the model consist of at least an impact word or phrase, such as &#8220;spannend&#8221; (<italic>suspenseful<xref ref-type="fn" rid="n3">3</xref></italic>). However, many rules require that there is also a book aspect term. For instance, the evaluative word &#8220;goed&#8221; (<italic>good</italic>) by itself can refer to anything. To be considered part of an impact expression, it must co-occur in one sentence with a word in one of the book aspect categories, e.g. a style-related word like &#8220;geschreven&#8221; (<italic>written</italic>) to be an expression of <italic>Aesthetic</italic> impact, or a narrative-related word like &#8220;verhaal&#8221; (<italic>story</italic>) or &#8220;plot&#8221; to be an expression of <italic>Narrative</italic> impact.</p>
<p>The DRIM identified 2,089,576 expressions of impact in the full review dataset. To identify the key impact terms per genre, we use the full review dataset with all of the approximately 2,1 Mio. impact expressions. To make a clearer distinction between impact expressions of generic affect and affect specific to narrative or aesthetics, we consider as <italic>Affect</italic> only those expressions that are not also categorized as <italic>Narrative</italic> or <italic>Aesthetic</italic>. Of the 2,089,576 expressions, there are 667,672 expressions for <italic>Aesthetic</italic> impact, 690,184 for <italic>Narrative</italic> impact and 731,720 for generic <italic>Affect</italic>.</p>
</sec>
<sec id="S3.2">
<title>3.2 Connecting Books and Reviews</title>
<p>A crucial step in relating topics in fiction to reading impact expressed in reviews is to connect the books to their corresponding reviews. For this, we rely mostly on the ISBN<xref ref-type="fn" rid="n4">4</xref> and the author and the book title. Note that a particular work may be connected to multiple ISBNs, for instance when reprints or new editions are produced for the same work with a different ISBN. Many mappings between reviews and books, and between multiple ISBNs of the same work were already made by Boot (<xref ref-type="bibr" rid="B8">2017</xref>) and Koolen et al. (<xref ref-type="bibr" rid="B22">2020</xref>), for the <italic>Online Dutch Book Response</italic> (ODBR) dataset of 472,810 reviews. We added around 160,000 reviews from Hebban to the ODBR set. To find ISBNs that refer to the same work, we first queried all ISBNs found in reviews using the SRU<xref ref-type="fn" rid="n5">5</xref> service of the National Library of the Netherlands. This SRU service gives access to the combined catalog of Dutch libraries and in many cases links multiple editions of the same work with different ISBNs. Using author and title, we resolved another number of duplicated works with different ISBNs. We then mapped all ISBNs of the same work to a unique work ID and linked the reviews via the ISBNs they mention to these work IDs. There are 125,542 distinct works reviewed by the reviews in our dataset. Of the 18,885 books for which we have ePubs, there are 10,056 books with at least one review in our data set. Altogether, these 10,056 unique works are linked to 130,751 reviews.</p>
</sec>
<sec id="S3.3">
<title>3.3 Connecting Impact and Topic Data</title>
<p>Our goal was to have a comprehensive mapping of the most relevant topics of works to their reviews, the latter analyzed via the DRIM. To create this dataset, we needed to connect the expressions of impact to the topics in our book dataset. To do so, we took the top five dominant topics of each book<xref ref-type="fn" rid="n6">6</xref> and linked those topics to the impact expressions in the reviews of the books for that topic. This resulted in a dataset in which each entry links specific reviews to the top five dominant topics for each book.</p>
<p>The Top2Vec model gave us a total of 228 topics. We attempted to label each topic with a distinct content label, but found that many topics are thematically very similar, capturing many of the same elements. Therefore, we manually assigned each topic to one or more of 19 broader themes: 1. <italic>geography &amp; setting</italic>, 2. <italic>behaviors/feelings</italic>, 3. <italic>culture</italic>, 4. <italic>crime</italic>, 5. <italic>history</italic>, 6. <italic>religion, spirituality &amp; philosophy</italic>, 7. <italic>supernatural, fantasy &amp; sci-fi</italic>, 8. <italic>war</italic>, 9. <italic>society</italic>, 10. <italic>city &amp; travel</italic>, 11. <italic>romance &amp; sex</italic>, 12. <italic>medicine/health</italic>, 13. <italic>wildlife/nature</italic>, 14. <italic>economy &amp; work</italic>, 15. <italic>lifestyle &amp; sport</italic>, 16. <italic>politics</italic>, 17. <italic>family</italic>, 18. <italic>science</italic>, 19. <italic>other</italic>. We provide the number of topics grouped per theme in <xref ref-type="fig" rid="F2">Figure 2</xref><xref ref-type="fn" rid="n7">7</xref>.</p>
<fig id="F2">
<caption>
<p><bold>Figure 2:</bold> The number of topics and books per theme.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g2.png"/>
</fig>
<p>We provide the full list of topics, themes, and their respective words in our code repository<xref ref-type="fn" rid="n8">8</xref>.</p>
</sec>
<sec id="S3.4">
<title>3.4 Book Genre Information</title>
<p>For genre information about books, we use the Dutch NUR<xref ref-type="fn" rid="n9">9</xref> classification codes assigned by publishers. As NUR was designed as a marketing instrument to determine where books are shelved in bookshops, publishers can choose codes based not only on the perceived genre of a book but also on marketing strategies related to where they want a book to be shelved to find the biggest audience. Some NUR codes refer to the same or very similar genres. E.g., codes 300, 301, and 302 refer to <italic>general literary fiction, Dutch literary fiction</italic>, and <italic>translated literary fiction</italic>, respectively, which we group together under <italic>Literary fiction</italic>. Similarly, we group codes 313, 330, 331, 332, and 339 under <italic>Suspense</italic> novels, as they all refer to types of suspense, i.e., <italic>pocket suspense, general suspense novels, detective novels</italic>, and <italic>thrillers</italic>, respectively. In total, we select 19 different NUR codes and map them to 9 genres. All remaining NUR codes in the fiction range (300-350) we map to <italic>Other fiction</italic> and the rest to <italic>Non-fiction</italic>. The full mapping is provided in <xref ref-type="sec" rid="A1">Appendix A</xref>.</p>
</sec>
<sec id="S3.5">
<title>3.5 Keyness Analysis on Impact Terms</title>
<p>The goal of this analysis is to determine (i) <italic>which</italic> words readers use in their reviews to describe the impact of a particular book and (ii) how <italic>characteristic</italic> these words are for a particular genre, compared to another genre. A good candidate to measure both (i) and (ii) is keyword analysis or keyness (<xref ref-type="bibr" rid="B16">Dunning 1994</xref>; <xref ref-type="bibr" rid="B18">Gabrielatos 2018</xref>; <xref ref-type="bibr" rid="B31">Paquot and Bestgen 2009</xref>).</p>
<p>There is ample literature comparing different keyness measures (<xref ref-type="bibr" rid="B12">Culpeper and Demmen 2015</xref>; <xref ref-type="bibr" rid="B15">Du et al. 2022</xref>; <xref ref-type="bibr" rid="B16">Dunning 1994</xref>; <xref ref-type="bibr" rid="B18">Gabrielatos 2018</xref>; <xref ref-type="bibr" rid="B25">Lijffijt et al. 2016</xref>) and finding that no single measure is perfect. A commonly used measure is <italic>G<sup>2</sup></italic>, which identifies <italic>key</italic> terms that occur statistically significantly more or less often in a target corpus (the reviews for a particular genre) compared to a reference corpus (reviews for one or more other genres).</p>
<p>Lijffijt et al. (<xref ref-type="bibr" rid="B25">2016</xref>) showed that Log-Likelihood Ratio (<italic>G<sup>2</sup></italic>, <xref ref-type="bibr" rid="B16">Dunning 1994</xref>) and several other frequency-based bag-of-words keyness measures suffer from excessively high confidence in their estimates because these measures assume samples to be statistically independent, but words in a text are not independent of each other. Du et al. (<xref ref-type="bibr" rid="B15">2022</xref>) compare frequency-based and dispersion-based measures for a downstream task (text classification) to show that for identifying key terms in a sub-corpus compared to the rest of the corpus, dispersion-based measures are more effective.</p>
<p>To compare the dispersion of a word or phrase in a target corpus to its dispersion in a reference corpus, Du et al. (<xref ref-type="bibr" rid="B14">2021</xref>) introduce <italic>Eta</italic>, which is a variant of the <italic>Zeta</italic> measure by Burrows (<xref ref-type="bibr" rid="B10">2006</xref>).</p>
<p>They find that <italic>Eta</italic> (<xref ref-type="bibr" rid="B14">Du et al. 2021</xref>) and <italic>Zeta</italic> (<xref ref-type="bibr" rid="B10">Burrows 2006</xref>) are among the most effective measures. Both <italic>Eta</italic> and <italic>Zeta</italic> compare document proportions of keywords. The former uses <italic>Deviation of Proportions</italic> (<italic>DP</italic>) (<xref ref-type="bibr" rid="B20">Gries 2008</xref>) which computes two sets of proportions. The first are the proportions that the lengths of documents represent with respect to the total number of words in a corpus (e.g., the set of reviews for books of a specific genre) as an expected distribution of the proportions of keywords. The second is the set of observed proportions of a keyword across a corpus with respect to the total corpus frequency of that keyword. There are two problems with using <italic>DP</italic> for keyness of impact terms. The first is that some impact terms do not occur in any of the reviews of a specific genre. In such cases, the observed proportions are not properly defined (a proportion of zero is not well-defined), so <italic>DP</italic> cannot be computed. The second is that the frequency distribution of impact terms in reviews is extremely skewed (84% of all impact terms in reviews have a frequency of 1, while 13% occur twice and the remaining 3% occur three or four times). Although longer reviews have a higher <italic>a priori</italic> probability of containing a specific impact term than shorter reviews, the frequency distribution of individual impact terms behaves more like a binomial distribution, so length-based proportions are not an appropriate measure of keyness.</p>
<p>Because of this, we instead measure dispersion using <italic>document frequencies</italic> (the number of reviews for a book genre in which an impact term occurs) to compute the <italic>document proportion</italic> (the fraction of reviews for a book genre in which an impact term occurs at least once). This gives the document proportion <inline-formula><mml:math id="Eq001-mml"><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>o</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>P</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> per impact term <italic>t</italic> and genre <italic>G</italic>, with the absolute difference <italic>Zeta</italic> between two genres defined as:</p>
<p><inline-formula><mml:math id="Eq002-mml"><mml:mrow><mml:mrow><mml:mi>Z</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>b</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>s</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>o</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>P</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#8722;</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>o</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>P</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula>.</p>
<p>To illustrate this approach, we compare the document proportions per genre of the impact terms &#8220;stijl&#8221; (<italic>style</italic>) and &#8220;schrijfstijl&#8221; (<italic>writing style</italic>). The former has the highest document proportion for reviews of <italic>Literary fiction</italic> (occurring in 3.7% of the reviews) and least in those of <italic>Non-fiction</italic> (1.2%), resulting in <inline-formula><mml:math id="Eq003-mml"><mml:mrow><mml:mrow><mml:mi>Z</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mn>0.037</mml:mn><mml:mo>&#8722;</mml:mo><mml:mn>0.012</mml:mn></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.025</mml:mn></mml:mrow></mml:math></inline-formula>. The latter is most common in reviews of <italic>Romance</italic> (14.6%) and least common in those of <italic>Non-fiction</italic> (2.0%), giving <inline-formula><mml:math id="Eq004-mml"><mml:mrow><mml:mrow><mml:mi>Z</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#8290;</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mn>0.146</mml:mn><mml:mo>&#8722;</mml:mo><mml:mn>0.02</mml:mn></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.126</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="S4">
<title>4. Results</title>
<sec id="S4.1">
<title>4.1 Topic and Genre</title>
<p>Zundert et al. (<xref ref-type="bibr" rid="B44">2022</xref>) found that the topics identified with Top2Vec are strongly associated with genre as identified by publishers. Similarly, Sobchuk and &#352;e&#316;a (<xref ref-type="bibr" rid="B37">2023</xref>) find that Doc2Vec &#8211; which is used by Top2Vec to embed the documents in the latent semantic space in which topic vectors are identified &#8211; is more effective at clustering books by genre than LDA (<xref ref-type="bibr" rid="B6">Blei et al. 2003</xref>).</p>
<sec id="S4.1.1">
<title>4.1.1 Genre Distribution per Topic</title>
<p>To extend the findings of Zundert et al. (<xref ref-type="bibr" rid="B44">2022</xref>), we first quantitatively demonstrate that there is a relationship between topic and genre. Each topic is associated with a number of books and thereby with the same number of genre labels. From eyeballing the distribution of genre labels per topic, it seems that for most topics, the vast majority of books in that topic belong to a single genre. But the genre distribution of the entire collection is also highly skewed, with a few very large genres and many much smaller genres. So perhaps the skew in most topics resembles the skew of the genre distribution of the collection.</p>
<p>To measure how much the genre distribution per topic deviates from that of the collection, we compute the Kullback-Leibler divergence (KL divergence) between the two distributions.<xref ref-type="fn" rid="n10">10</xref> This gives a set of 228 deviations from the collection distribution.</p>
<p>But whether these deviations are small or large is difficult to read from the numbers themselves. For that, we should compare them against a random shuffling of the book genres across books (while keeping the books assigned per topic stable). For large topics (with many books), a random shuffling should have a genre distribution close to that of the collection. For small clusters, the divergence will tend to be higher.</p>
<p>We create five alternative clusterings with books randomly assigned to topics with the same topic size distribution as established by the topic model. The distribution of the 228 KL divergence scores per model (five random and one topic model) are shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. The five random models have almost identical distributions concentrated around 0.1 with a standard deviation of around 0.075 and a max. of around 0.5. The genre distribution of the topic model is very different, with a median score of 1.06 and more than 75% of all scores above 0.68. From this quantitative analysis, it is clear that there is a strong relationship between topic and genre.</p>
<fig id="F3">
<caption>
<p><bold>Figure 3:</bold> The KL divergence between the genre distribution per topic and that of the collection for the topic model as well as for five random shuffles of the genre labels using the same books per topic.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g3.png"/>
</fig>
<p>We can use the same random shuffling to get more insight into how topics cluster genres. For that, we compute the <italic>observed</italic> co-occurrence of pairs of genres by iterating over all pairs of books in each topic and counting the co-occurence of their respective genres and divide that by the <italic>expected</italic> co-occurrence of pairs of genres when the books are randomly shuffled. For the <italic>expected</italic> co-occurrence, each shuffling gives different counts, so we repeat the random shuffling 100 times and take the mean number of co-occurrences per pair of genres as the expectation. The Observed over Expected (OoE) ratio is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. An OoE ratio of 1 means that two genres are co-occuring no more in the topics than is expected when there is no relationship between topic and genre. Scores higher than 1 mean genres are more likely to co-occur than chance (topically, they are similar to each other) and lower than 1 that they are <italic>less</italic> likely to co-occur (topically, they are dissimilar to each other). The numbers on the diagonal are the highest per row and column, meaning that books of each genre are more likely to end up in topics with other books of the same genre than with books of a different genre.</p>
<fig id="F4">
<caption>
<p><bold>Figure 4:</bold> Observed over Expected ratio (OoE) of genre co-occurrences as observed in the 228 topics compared to the expected co-occurrences of randomly shuffling the books over 228 clusters of the same size.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g4.png"/>
</fig>
<p>We make a few more observations. First, some genres are very dissimilar from others. Most of the compared OoE scores for <italic>Fantasy</italic> with other genres are well below 1.0. It is topically only slightly similar to <italic>Young adult</italic>. Second, some genres are topically similar to each other. <italic>Children&#8217;s fiction</italic> and <italic>Young adult</italic> have an OoE of 3.52, while the OoE of <italic>Literary thriller</italic> and <italic>Suspense</italic> is 2.17. These are topical connections that are not surprising. Third, <italic>Literary fiction</italic> is topically somewhat similar to <italic>Literary thrillers</italic> (OoE of 1.23), but dissimilar to <italic>Suspense</italic> (0.46). Even though <italic>Literary thriller</italic> is similar to <italic>Suspense</italic>, it has a topical connection to other <italic>Literary fiction</italic> that <italic>Suspense</italic> does not have. In addition, while NUR codes are mostly a marketing instrument, their distinction between <italic>Literary thrillers</italic> and other <italic>Suspense</italic> novels relates to a topical distinction as well. Finally, fourth, the numbers on the diagonal vary strongly, with <italic>Historical fiction</italic> novels being much more likely to be topically clustered with other historical novels than with novels of other genres (OoE is 24.23), while for <italic>Literary fiction</italic> (2.45) and <italic>Non-fiction</italic> (3.23) this is much less likely. This may be partly due to the fact that the latter two are the largest genres in the collection and therefore have a high <italic>a priori</italic> probability to end up in topics with books of other genres, but we speculate that it may also be due to the fact that these two genres do not have a clear topic profile (whereby we stress that topic here is interpreted as sharing vocabulary, because the Doc2Vec embedding space is based on word tokens).</p>
</sec>
<sec id="S4.1.2">
<title>4.1.2 Thematic Distribution per Genre</title>
<p>Next, we perform a qualitative analysis of the topics and their relationship to genre, via the identified themes described in <xref ref-type="sec" rid="S3.3">subsection 3.3</xref>.</p>
<p>The distribution of topic themes per genre is shown in <xref ref-type="fig" rid="F5">Figure 5</xref> in the form of radar plots. The genres show distinct thematic profiles. <italic>Literary fiction</italic> scores high on the themes of <italic>culture, geography &amp; setting</italic> and <italic>behaviors/feelings</italic> which is perhaps not surprising. <italic>Non-fiction</italic> scores high on <italic>religion, spirituality &amp; philosophy, medicine/health, economy &amp; work</italic>, and <italic>behaviors/feelings</italic> which are themes that few fiction genres score high on.</p>
<fig id="F5">
<caption>
<p><bold>Figure 5:</bold> Radar plots showing the relative prevalence of themes in six genres, from left to right, top to bottom: <italic>Literary thrillers, Suspense, Children&#8217;s fiction</italic> and <italic>Young adult, Romance, Fantasy, Literary fiction, Historical fiction, Other fiction</italic> and <italic>Non-fiction</italic>.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g5.png"/>
</fig>
<p>In <italic>Children&#8217;s fiction</italic>, there is relatively little use of the geographical aspect of setting, especially compared to other fiction genres. That is, it seems that children&#8217;s novels make little explicit reference to geographical places. They score high on <italic>behaviors/feelings</italic> and moderately high on <italic>culture, family</italic> and <italic>supernatural, fantasy &amp; sci-fi</italic>. The main difference between <italic>Children&#8217;s fiction</italic> and <italic>Young adult</italic> is that the latter scores higher on <italic>supernatural, fantasy &amp; sci-fi</italic>. For the former, <italic>Young adult</italic> strongly overlaps with <italic>Fantasy</italic> novels. <italic>Young adult</italic> also adds in a bit of <italic>romance &amp; sex</italic>. These observations suggest that <italic>Children&#8217;s fiction</italic> and <italic>Young adult</italic> by and large treat the same themes, but against different &#8216;backgrounds&#8217;. <italic>Children&#8217;s fiction</italic> deals with <italic>behaviors/feelings</italic> against a backdrop of <italic>culture</italic> and <italic>family</italic>. <italic>Young adult</italic> does practically the same, but adds <italic>supernatural, fantasy &amp; sci-fi</italic> elements to the story and opens the stage for some romantic behavior.</p>
<p>If one were to hazard a guess about reader development, it would almost seem as if young readers are invited to pre-sort on the major themes of grown-up literature, whith <italic>Romance</italic> amplifying the <italic>romance &amp; sex</italic> encountered in <italic>Young adult</italic> books, while <italic>Literary fiction</italic> and <italic>Literary thrillers</italic> amplify motifs of <italic>culture, setting</italic>, and <italic>crime</italic>, and <italic>Fantasy</italic> caters to the interest in the supernatural developed through <italic>Young adult</italic> fiction. Much more research would be needed, however, to substantiate such a pre-sorting effect. In any case, <italic>Romance</italic> scores high on <italic>romance &amp; sex</italic> and has medium scores for <italic>culture</italic> and <italic>geography &amp; setting</italic>, while <italic>Suspense</italic> novels score high on <italic>crime</italic> and have medium scores for <italic>geography &amp; setting</italic> and <italic>war</italic>.</p>
<p>We expect that many of these observations coincide with intuitions of literary researchers. This suggests that the grouping of topics by theme makes sense from a literary analytical perspective. The findings also show where genres overlap and where they differ. For instance, the profile for <italic>Literary fiction</italic> and <italic>Literary thriller</italic> are similar, with the main difference being the much higher prevalence of the <italic>crime</italic> theme in <italic>Literary thrillers</italic>. <italic>Suspense</italic> is similar to <italic>Literary thrillers</italic> in the prevalence of <italic>crime</italic> as theme, but lower scores for <italic>culture</italic> and <italic>geography &amp; setting</italic>.</p>
<p>One of the main findings is that, for the chosen document frequency range of mid-frequency terms, there is a clear connection between topic and genre, with thematic clustering of topics leading to distinct genre profiles, but also to thematic connections between certain genres. None of this will radically transform our understanding of genre and topic, but it prompts the question how different parts of the document frequency distribution relate to different aspects of novels. From authorship attribution research, we know that authorial signal is mainly found in the high-frequency range and our work corroborates earlier findings that topics contain genre-signals in mid-range frequencies (<xref ref-type="bibr" rid="B38">Thompson and Mimno 2018</xref>; <xref ref-type="bibr" rid="B44">Zundert et al. 2022</xref>).</p>
</sec>
</sec>
<sec id="S4.2">
<title>4.2 Impact and Genre</title>
<sec id="S4.2.1">
<title>4.2.1 Reviews per Genre</title>
<p>With the genre labels, we can count how many books in each genre have reviews in our dataset and how many reviews they have (<xref ref-type="table" rid="T1">Table 1</xref>). The genre with the highest total number of reviews is <italic>Literary fiction</italic>, with 200,907 reviews in our dataset, followed by <italic>Literary thrillers</italic> and <italic>Suspense</italic> novels. If we consider the number of reviews per book, <italic>Literary thrillers</italic> have the highest mean number of reviews (22.8). However, the distribution of the number of reviews per book is highly skewed, with a single review per book being the most likely and having more reviews being increasingly unlikely (<xref ref-type="bibr" rid="B22">Koolen et al. 2020</xref>). The distributions per genre show some differences, but all are close to a power-law. The cumulative distribution function of the number of reviews per book for the different genres are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, with on the Y-axis the probability <inline-formula><mml:math id="Eq006-mml"><mml:mrow><mml:mi>P</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>&#8805;</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> that a book has at least <italic>x</italic> reviews.<xref ref-type="fn" rid="n11">11</xref></p>
<table-wrap id="T1">
<caption>
<p><bold>Table 1:</bold> Reviews per genre and mean number of reviews per book per genre.</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"></td>
<td align="right" valign="top">Reviewed books</td>
<td align="right" valign="top">Reviews</td>
<td align="right" valign="top">Mean reviews/book</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Literary fiction</td>
<td align="right" valign="top">19,288</td>
<td align="right" valign="top">200,907</td>
<td align="right" valign="top">10.4</td>
</tr>
<tr>
<td align="left" valign="top">Literary thriller</td>
<td align="right" valign="top">3,394</td>
<td align="right" valign="top">77,288</td>
<td align="right" valign="top">22.8</td>
</tr>
<tr>
<td align="left" valign="top">Young adult</td>
<td align="right" valign="top">2,919</td>
<td align="right" valign="top">30,552</td>
<td align="right" valign="top">10.5</td>
</tr>
<tr>
<td align="left" valign="top">Children fiction</td>
<td align="right" valign="top">5,348</td>
<td align="right" valign="top">27,989</td>
<td align="right" valign="top">5.2</td>
</tr>
<tr>
<td align="left" valign="top">Suspense</td>
<td align="right" valign="top">6,266</td>
<td align="right" valign="top">67,990</td>
<td align="right" valign="top">10.9</td>
</tr>
<tr>
<td align="left" valign="top">Fantasy fiction</td>
<td align="right" valign="top">1,571</td>
<td align="right" valign="top">13,739</td>
<td align="right" valign="top">8.7</td>
</tr>
<tr>
<td align="left" valign="top">Romance</td>
<td align="right" valign="top">1,291</td>
<td align="right" valign="top">6,434</td>
<td align="right" valign="top">5.0</td>
</tr>
<tr>
<td align="left" valign="top">Historical fiction</td>
<td align="right" valign="top">556</td>
<td align="right" valign="top">3,463</td>
<td align="right" valign="top">6.2</td>
</tr>
<tr>
<td align="left" valign="top">Regional fiction</td>
<td align="right" valign="top">472</td>
<td align="right" valign="top">1,528</td>
<td align="right" valign="top">3.2</td>
</tr>
<tr>
<td align="left" valign="top">Other fiction</td>
<td align="right" valign="top">7,260</td>
<td align="right" valign="top">37,515</td>
<td align="right" valign="top">5.2</td>
</tr>
<tr>
<td align="left" valign="top">Non-fiction</td>
<td align="right" valign="top">26,884</td>
<td align="right" valign="top">109,158</td>
<td align="right" valign="top">4.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6">
<caption>
<p><bold>Figure 6:</bold> The cumulative distribution function of the number of reviews per book, on a log-log scale. The Y-axis shows the probability <inline-formula><mml:math id="Eq005-mml"><mml:mrow><mml:mi>P</mml:mi><mml:mo>&#8290;</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>&#8805;</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> that a book has at least <italic>x</italic> reviews.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g6.png"/>
</fig>
<p>The curves for some of the genres overlap, which makes them difficult to discern, but there are a few main insights. First, <italic>Regional fiction</italic> and <italic>Non-fiction</italic> have the fastest falling curves, indicating that books in these genres are the least likely to acquire many reviews. Next is a cluster of <italic>Children&#8217;s fiction, Romance, Historical fiction</italic>, and <italic>Other fiction</italic>, which tend to get a slightly higher number of reviews. Then there is a cluster of <italic>Suspense, Literary fiction, Young adult</italic>, and <italic>Fantasy fiction</italic>, which tend to get more reviews than the previous cluster. And finally, clearly above the rest, is the curve of <italic>Literary thrillers</italic>, which tend to get more reviews than books in any other genre.</p>
<p><italic>Thrillers</italic> are more often reviewed on the platforms that are in the review dataset. <italic>Romance</italic> novels have fewer reviews but are a very popular genre (Regis (<xref ref-type="bibr" rid="B35">2003, 108</xref>), see also Darbyshire (<xref ref-type="bibr" rid="B13">2023</xref>)). This prompts the question of whether readers of <italic>Regional</italic> and <italic>Romance</italic> novels have less desire to review these novels or review them on different platforms and in different ways. As there seem to be many video reviews of <italic>Romance</italic> novels on TikTok using the tag #BookTok, this would be a valuable resource to add to our investigations. A difference in the number of reviews might be a signal of a difference in impact, but it is also plausible that different genres attract different types of readers who express their impact in different ways linguistically, using different media (e.g., text or video) on different platforms (e.g., GoodReads or TikTok). To that extent, the review dataset may be a biased representation of the impact of books in different genres. Bracketing for a moment the potential skewedness of the number of reviews per genre and taking the number of reviews as a proxy of popularity, it is also interesting to observe that popularity is apparently a commodity that is reaped in orders of magnitude.</p>
</sec>
<sec id="S4.2.2">
<title>4.2.2 Key Impact Terms per Genre</title>
<p><bold>Correlations between genres</bold> First, we compare genres in terms of their impact terms using the document proportions per impact term. For each pair of genres, we compute the Pearson correlation <inline-formula><mml:math id="Eq007-mml"><mml:mi>&#961;</mml:mi></mml:math></inline-formula> between the document proportions of all impact terms. A high positive correlation means that impact terms with a high document proportion in one genre tend to also have a high document proportion in the other genre.</p>
<p>The correlations per impact type are shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. For <italic>Affect</italic> impact terms (the top correlation table), most genre pairs have a near perfect correlation (<inline-formula><mml:math id="Eq008-mml"><mml:mrow><mml:mn>0.8</mml:mn><mml:mo>&lt;</mml:mo><mml:mi>&#961;</mml:mi><mml:mo>&lt;</mml:mo><mml:mn>1.0</mml:mn></mml:mrow></mml:math></inline-formula>) and only few pairs have a moderate (<inline-formula><mml:math id="Eq009-mml"><mml:mrow><mml:mn>0.4</mml:mn><mml:mo>&lt;</mml:mo><mml:mi>&#961;</mml:mi><mml:mo>&lt;</mml:mo><mml:mn>0.6</mml:mn></mml:mrow></mml:math></inline-formula>) or strong correlation (<inline-formula><mml:math id="Eq010-mml"><mml:mrow><mml:mn>0.6</mml:mn><mml:mo>&lt;</mml:mo><mml:mi>&#961;</mml:mi><mml:mo>&lt;</mml:mo><mml:mn>0.8</mml:mn></mml:mrow></mml:math></inline-formula>), notably <italic>Children&#8217;s fiction</italic> in combination with either <italic>Historical ficton, Literary thrillers</italic> and <italic>Suspense</italic>. For <italic>Narrative</italic> impact terms, there are more moderate correlations, with <italic>Non-fiction</italic> standing out as the most distinct genre. This is not surprising, given that (we assume) <italic>Non-fiction</italic> books are least likely to be discussed in terms of narrative. For <italic>Aesthetic</italic> impact terms, there are only four correlations below but close to 0.8, indicating that there are few differences in vocabulary between genres. The overwhelming majority of strong and near perfect correlations suggests that, overall, impact across genres is expressed in the same vocabulary.</p>
<fig id="F7">
<caption>
<p><bold>Figure 7:</bold> Pearson correlation in the doc proportion scores of impact terms between pairs of genres, for <italic>Affect</italic> (top), <italic>Narrative</italic> (middle) and <italic>Aesthetic</italic> (bottom).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g7.png"/>
</fig>
<p><bold>Vocabulary differences between genres</bold> Even though the correlations are mostly strong, we can still zoom in on the largest differences in vocabulary usage. For generic <italic>Affect, Children&#8217;s fiction</italic> is most distinctive as it has high score differences with all other genres. The document proportions for generic <italic>Affect</italic> terms of <italic>Children&#8217;s fiction</italic> and <italic>Regional fiction</italic> are shown in <xref ref-type="fig" rid="F8">Figure 8</xref>. The diagonal line shows where terms have equal proportions in both genres. Reviews of <italic>Children&#8217;s fiction</italic> seem to use a smaller impact vocabulary &#8211; almost all document proportions are close to zero &#8211; but much higher proportions for the impact term &#8220;leuk&#8221; (<italic>fun</italic> or <italic>cool</italic>). This term is used much less in reviews of other genres.</p>
<fig id="F8">
<caption>
<p><bold>Figure 8:</bold> Document proportions of generic <italic>Affect</italic> terms for <italic>Children&#8217;s fiction</italic> and <italic>Regional fiction</italic>.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g8.png"/>
</fig>
<p>For <italic>Narrative</italic> impact, the biggest summed difference is between <italic>Romance</italic> and <italic>Literary thrillers</italic> (see <xref ref-type="fig" rid="F9">Figure 9</xref>). The main differences are found with a handful of terms, &#8220;spannend&#8221; (<italic>thrilling, suspenseful</italic>), &#8220;spanning&#8221; (<italic>suspense</italic>) and &#8220;verrassen&#8221; (<italic>surprise</italic>) are more common in <italic>Literary thrillers</italic> and &#8220;romantisch&#8221; (<italic>romantic</italic>) and &#8220;heerlijk&#8221; (<italic>lovely, wonderful</italic>) are more common in <italic>Romance</italic> novels. These are perhaps somewhat obvious, but show that impact, or at least the language of impact, is related to genre.</p>
<fig id="F9">
<caption>
<p><bold>Figure 9:</bold> Document proportions of <italic>Narrative</italic> impact terms for <italic>Romance</italic> and <italic>Literary thrillers</italic>.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g9.png"/>
</fig>
<p>For <italic>Aesthetic</italic> impact, the biggest summed difference is between <italic>Romance</italic> and <italic>Historical fiction</italic> (see <xref ref-type="fig" rid="F10">Figure 10</xref>). Here again, the main differences are in a few terms. Reviews of <italic>Historical fiction</italic> more often mention impact terms like &#8220;mooi&#8221; (<italic>beautiful</italic>), &#8220;beschrijven&#8221; (<italic>describe</italic>), &#8220;beschreven&#8221; (<italic>described</italic>), and &#8220;prachtig&#8221; (<italic>beautiful</italic>). Reviews of <italic>Romance</italic> novels more often mention &#8220;schrijfstijl&#8221; (<italic>writing style</italic>), &#8220;humor&#8221; (<italic>humor</italic>), and &#8220;luchtig&#8221; (<italic>airy</italic>). It seems that for <italic>Historical fiction</italic>, reviewers focus more on descriptions (how evocatively the author describes historical settings, persons, or events), while reviewers of <italic>Romance</italic> novels focus more on humor and lightness of style. A close reading of some of the contexts in which &#8220;schrijfstijl&#8221; is mentioned in <italic>Romance</italic> reviews suggests that reviewers often use it in phrases like &#8220;makkelijke schrijfstijl&#8221; and &#8220;vlotte schrijfstijl&#8221; (<italic>a writing style that reads easily</italic> or <italic>quickly</italic>, respectively).</p>
<fig id="F10">
<caption>
<p><bold>Figure 10:</bold> Document proportions of <italic>Aesthetic</italic> impact terms for <italic>Historical fiction</italic> and <italic>Romance</italic>.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g10.png"/>
</fig>
</sec>
</sec>
<sec id="S4.3">
<title>4.3 Impact and Topic</title>
<p>The third link between the three main concepts that are the focus of this paper is between impact and topic.</p>
<p>To study how the use of impact terms differs between reviews of books with different themes &#8211; recall, we are talking about theme in the sense of topically grouped clusters of books &#8211; we first need to group the reviews by theme. Because themes are based on topics and some themes share the same topics, some reviews are assigned to multiple themes. We calculated Pearson correlations between themes in terms of the document proportions per impact term, just as we did for genre (see <xref ref-type="fig" rid="F13">Figure 13</xref>, <xref ref-type="fig" rid="F14">Figure 14</xref>, and <xref ref-type="fig" rid="F15">Figure 15</xref> in <xref ref-type="sec" rid="A3">Appendix C</xref>). There are many observations that could be made, but again, we limit ourselves to the most salient ones related to the three largest themes (in number of books). First of all, the vast majority of the correlations are near perfect, suggesting that impact is expressed with similar vocabulary across reviews for books associated with different themes. For <italic>Aesthetic</italic> impact, there are no correlations below 0.8. For <italic>Narrative</italic> impact, the one clearly distinct theme is <italic>medicine/health</italic>, which has no or weak correlations with most of the other themes.</p>
<p>When we zoom in on the document proportions of individual impact terms and compare two genres, we observe the overall similarity but also some specific differences. The comparative document proportions for general <italic>affect</italic> terms are shown for reviews of books related to themes <italic>crime</italic> and <italic>culture</italic> (top of <xref ref-type="fig" rid="F11">Figure 11</xref>) or <italic>family</italic> and <italic>war</italic> (bottom of <xref ref-type="fig" rid="F11">Figure 11</xref>).</p>
<fig id="F11">
<caption>
<p><bold>Figure 11:</bold> Document proportions of general <italic>Affect</italic> terms for the themes <italic>crime</italic> and <italic>culture</italic> (top) and <italic>family</italic> and <italic>war</italic> (bottom).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g11.png"/>
</fig>
<p>The proportions for <italic>crime</italic> and <italic>culture</italic> are slightly different, but most of the data points are close to the diagonal and the correlation between the sets of proportions is high. Impact terms like &#8220;verrasen&#8221; (<italic>to surprise</italic>) and &#8220;aanrader&#8221; (<italic>recommendation</italic>) are used slightly more often in reviews of <italic>crime</italic>-related books, while terms like &#8220;gevoel&#8221; (<italic>feeling</italic>), &#8220;emotie&#8221; (<italic>emotion</italic>), and &#8220;grappig&#8221; (<italic>funny</italic>) are more often used for <italic>culture</italic>-related books. However, the relative differences in proportion are small. For <italic>family</italic> and <italic>history</italic>, we observe larger differences, with <italic>affect</italic> terms like &#8220;leuk&#8221; (<italic>fun, enjoyable</italic>) and especially &#8220;grappig&#8221; (<italic>funny</italic>) having much higher document proportions in reviews of <italic>family</italic>-related books than <italic>war</italic>-related books.</p>
<p>For <italic>Narrative</italic> impact terms, the comparative document proportions for reviews related to the themes <italic>family</italic> and <italic>history</italic> are shown in <xref ref-type="fig" rid="F12">Figure 12</xref>. Again, most of the data points are close to the diagonal, showing the similarity in usage of impact terms. However, the biggest difference is that <italic>family</italic> reviewers are more likely to use terms like &#8220;herkenbaar&#8221; (<italic>recognizable</italic>), &#8220;ontroerend&#8221; (<italic>touching</italic>), while <italic>history</italic> reviewers more often use &#8220;indrukwekkend&#8221; (<italic>impressive</italic>), &#8220;aangrijpend&#8221; (<italic>gripping</italic>), and &#8220;boeiend&#8221; (<italic>intriguing, fascinating</italic>).</p>
<fig id="F12">
<caption>
<p><bold>Figure 12:</bold> Document proportions of <italic>Narrative</italic> impact terms for the themes <italic>family</italic> and <italic>history</italic>.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g12.png"/>
</fig>
<p>Note that for terms with lower document proportions (i.e., between 0 and 0.1%), the relative differences in proportions can be large, signaling potentially highly statistically significant differences between genres or themes. But, the fact that the proportions are low means that these significant differences are between very rare and extremely rare usage of terms. To illustrate, the <italic>Aesthetic</italic> term &#8220;geniaal&#8221; (<italic>genius</italic> or <italic>brilliant</italic>) is six times more likely to occur in <italic>Young adult</italic> fiction reviews than in reviews of <italic>Children&#8217;s fiction</italic>, but &#8220;geniaal&#8221; is very rare in the former (14 total occurrences, or 0.05% of 29,075 reviews) and extremely rare in the latter (two occurrences, or 0.008% of 25,074 reviews).</p>
<p>Although such large relative differences may give further insight into how genre and theme relate to impact, we want to stress the high overall similarity. It suggests that reviewers use a largely common vocabulary for expressing impact, regardless of the genre or theme of a book. Large relative differences in rare terms are potentially insightful to interpret differences between individual books, authors, or reviewers, but they say little about genres or topics overall.</p>
</sec>
</sec>
<sec id="S5">
<title>5. Discussion and Conclusion</title>
<p>In this paper, we investigated the relationship between three important concepts in literary studies: genre, topic, and impact. We discuss our findings for each pair of concepts in turn.</p>
<p><bold>Genre and topic</bold> Our analyses have corroborated earlier findings on the relationship between genre and topic. By clustering topics identified by topic modeling into broader themes and by measuring the prevalence of these themes in the books of specific genres, we find that topics have a strong relation with genres and the genres have distinct thematic profiles. These profiles match existing intuitions about the distribution of themes across genres. Potentially, these profiles can provide additional insight into genre dynamics (e.g., as to what motivates readers to mix-read genres or not), although much of this aspect remains to be examined.</p>
<p><bold>Genre and impact</bold> The Dutch Reading Impact Model (DRIM by Boot and Koolen (<xref ref-type="bibr" rid="B9">2020</xref>)) identifies sets of words that are to some extent related to genre, and by studying the overlap in key impact terms between genres, we find clusters of genres that are similar in how their impact is described. Of course, this is not entirely surprising. For instance, <italic>Suspense</italic> novels and <italic>Literary thrillers</italic> are highly similar in terms of all three types of impact. However, it is much less obvious or intuitive that <italic>Historical fiction, Literary fiction</italic> and <italic>Fantasy</italic> have very similar distributions of <italic>Aesthetic</italic> impact terms, nor that <italic>Non-fiction</italic> is distinct from most other genres in terms of <italic>Narrative</italic> impact, apart from <italic>Literary fiction</italic> and <italic>Other fiction</italic>.</p>
<p>It remains unclear for now how we should explain the relationship between impact and genre. Perhaps this relation signals that reviewers develop and copy conventions for writing about books from other reviews they have read, regardless of genre differences. At the same time, we should not ignore the differences that do exist. At an aggregate level, differences may seem small, but small differences in usage across a range of impact terms could still signal a consistent and meaningful difference in impact. Finally, depending on how the reading impact model was developed, this may also be an artifact of how the rules were constructed. For instance, if reviews for a heterogeneous set of books were scanned to identify recurring expressions of impact, it is possible that expressions that are shared across genres stood out and were more likely to be included in the set of rules. Further analysis is required to establish which, if any, of these factors contributes to the relationship between fiction genres and reading impact as expressed in reviews.</p>
<p><bold>Topic and impact</bold> For the first two pairs of concepts, there were some expectations, e.g., that there is a relation between the <italic>Romance</italic> genre and topics related to the theme of <italic>romance &amp; sex</italic>, or that typical narrative impact terms in reviews of <italic>Young adult</italic> novels overlap with those in reviews of <italic>Fantasy</italic> novels. For the link between topic and impact, we struggled to come up in advance with expectations on how the topics in novels are related to impact. Novels discussing topics such as <italic>war</italic> and its consequences or living with physical or mental illness might lead to more reviews mentioning <italic>Narrative</italic> impact. But honest reflection forces us to admit that the results of topic modeling do not shed much light on how authors deal with topics and how reviewers discuss them. This gap stubbornly persists throughout continued engagement with our data in several papers. Consequently, this should give us pause to reflect on our operationalizations. Although vector models have moved beyond bag-of-words approaches and are becoming increasingly more sophisticated, we have not inched significantly closer to answering the question of which features of novel texts relate to what types of reader impact adequately and satisfyingly from a literary studies perspective.</p>
<p>Our reflections tie in with observations and suggestions made in some recent methodological publications on computational humanities: Bode (<xref ref-type="bibr" rid="B7">2023</xref>) argues that humanities researchers applying conventional methods and those embracing computational or data science methods should take a greater and more sincere interest in each other&#8217;s work. Rather than addressing research questions by stretching either method beyond its limits, researchers ought to investigate how the different methods can reinforce and amplify each other. Pichler and Reiter (<xref ref-type="bibr" rid="B32">2022</xref>) argue that operationalizations in computational linguistics and computational literary studies are currently often poor because we typically fail to express the precise operations that identify the theoretical concept we are trying to observe. Indeed our operationalizations seem underwhelming in the light of literary mechanisms. The reason to label a topic as being <italic>about</italic> war is that it contains words directly and strongly associated with war and emphasizing the physical aspects of it, such as <italic>war, soldier, bombing, battlefield, wounded</italic>, etc. But novels that readers would describe as being <italic>about</italic> war might instead focus on more indirect aspects or on aspects that war shares with many other situations, such as dire living conditions or being cut-off from the rest of the world, feeling unsafe and scared, or the sense of helplessness or hopelessness. The problem is not just that words indirectly related to war might lead an annotator to label a topic as being about something other than war. It is also that an author, going by the good practice of &#8220;show don&#8217;t tell&#8221;, can conjure up images that fit these words in almost infinitely many ways that are almost impossible to capture by looking at bags of words. Which means we need infinitely better operationalizations.</p>
</sec>
<sec id="S6">
<title>6. Data Availability</title>
<p>Data used for the research can be found at: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact">https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact</ext-link>. It has been archived and is persistently available at: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.13929510">https://doi.org/10.5281/zenodo.13929510</ext-link>.</p>
</sec>
<sec id="S7">
<title>7. Software Availability</title>
<p>All code created and used in this research has been published at: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact">https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact</ext-link>. It has been archived and is persistently available at: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.13929510">https://doi.org/10.5281/zenodo.13929510</ext-link>.</p>
</sec>
</body>
<back>
<fn-group>
<fn id="n1"><p>See: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://www.hebban.nl/">https://www.hebban.nl/</ext-link>.</p></fn>
<fn id="n2"><p>Experiments with using different frequency ranges for filtering suggests that the topic modeling process is relatively insensitive with regards to the upper limit. I.e., using 50%, 30%, or 10% results in roughly equal numbers of topics that show the same relationship with book genre (see <xref ref-type="sec" rid="S4.1.1">subsubsection 4.1.1</xref> and the following notebook: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact/blob/main/notebooks/topic_and_genre.ipynb">https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact/blob/main/notebooks/topic_and_genre.ipynb</ext-link>.</p></fn>
<fn id="n3"><p>For all Dutch terms we will consistently provide English translation in italics between parentheses.</p></fn>
<fn id="n4"><p>International Standard Book Number, see: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://en.wikipedia.org/wiki/ISBN">https://en.wikipedia.org/wiki/ISBN</ext-link>.</p></fn>
<fn id="n5"><p>Search and Retrieval by URL, see: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://en.wikipedia.org/wiki/Search/Retrieve_via_URL">https://en.wikipedia.org/wiki/Search/Retrieve_via_URL</ext-link>.</p></fn>
<fn id="n6"><p>Top2Vec creates topics by clustering the document vectors and taking the centroid of each cluster as the topic vector. We computed the cosine similarity between the document vector (representing the book) and the topic vectors, and selected the top five closest (i.e., most similar) topics to each book.</p></fn>
<fn id="n7"><p>Note that in this paper &#8220;theme&#8221; should not be taken to coincide with the literary studies sense of theme. Rather we use the term &#8220;theme&#8221; to clearly distinguish between the topics as identified by Top2Vec and their clustering as done by us.</p></fn>
<fn id="n8"><p>See: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact/blob/main/data/topic_labels.tsv">https://github.com/impact-and-fiction/jcls-2024-topic-genre-impact/blob/main/data/topic_labels.tsv</ext-link>.</p></fn>
<fn id="n9"><p>NUR stands for Nederlandse Uniforme Rubrieksindeling or Dutch Uniform Categories classification.</p></fn>
<fn id="n10"><p>The KL divergence measures the statistical distance between two distributions, that is, how statistically different they are with respect to each other.</p></fn>
<fn id="n11"><p>We show the cumulative distribution instead of the plain distribution because it produces smoother curves and better shows the trends.</p></fn>
</fn-group>
<sec id="S8">
<title>8. Acknowledgements</title>
<p>This project has been supported through generous material and in-kind technical and data-science analytical support from the eScience Center in Amsterdam. We thank the National Library of the Netherlands for providing access to the novels used in this research and for their invaluable technical support.</p>
</sec>
<sec id="S9">
<title>9. Author Contributions</title>
<p><bold>Marijn Koolen:</bold> Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Writing &#8211; original draft</p>
<p><bold>Joris J. van Zundert:</bold> Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Visualization, Writing &#8211; review &amp; editing</p>
<p><bold>Eva Viviani:</bold> Formal analysis, Software, Validation, Visualization</p>
<p><bold>Carsten Schnober:</bold> Resources, Software</p>
<p><bold>Willem van Hage:</bold> Methodology, Resources, Software</p>
<p><bold>Katja Tereshko:</bold> Writing &#8211; original draft, Writing &#8211; review &amp; editing</p>
</sec>
<ref-list>
<ref id="B1"><mixed-citation publication-type="book"><string-name><surname>Abrams</surname>, <given-names>Meyer H.</given-names></string-name> (<year>1971</year>). <source>The Mirror and the Lamp: Romantic Theory and the Critical Tradition</source>. <publisher-name>Oxford University Press</publisher-name>.</mixed-citation></ref>
<ref id="B2"><mixed-citation publication-type="journal"><string-name><surname>Angelov</surname>, <given-names>Dimo</given-names></string-name> (<year>2020</year>). <article-title>&#8220;Top2Vec: Distributed Representations of Topics&#8221;</article-title>. In: <source>arXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2008.09470</pub-id>.</mixed-citation></ref>
<ref id="B3"><mixed-citation publication-type="webpage"><string-name><surname>Ashok</surname>, <given-names>Vikas Ganjigunte</given-names></string-name>, <string-name><given-names>Song</given-names> <surname>Feng</surname></string-name>, and <string-name><given-names>Yejin</given-names> <surname>Choi</surname></string-name> (<year>2013</year>). <article-title>&#8220;Success with Style: Using Writing Style to Predict the Success of Novels&#8221;</article-title>. In: <source>Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>, <fpage>1753</fpage>&#8211;<lpage>1764</lpage>. <uri>https://api.semanticscholar.org/CorpusID:7100691</uri> (visited on 07/28/2023).</mixed-citation></ref>
<ref id="B4"><mixed-citation publication-type="journal"><string-name><surname>Berger</surname>, <given-names>Jonah</given-names></string-name> and <string-name><given-names>Grant</given-names> <surname>Packard</surname></string-name> (<year>2018</year>). <article-title>&#8220;Are Atypical Things More Popular?&#8221;</article-title> In: <source>Psychological Science</source> <volume>29</volume> (<issue>7</issue>), <fpage>1178</fpage>&#8211;<lpage>1184</lpage>. <pub-id pub-id-type="doi">10.1177/0956797618759465</pub-id>.</mixed-citation></ref>
<ref id="B5"><mixed-citation publication-type="book"><string-name><surname>Berry</surname>, <given-names>David M.</given-names></string-name> (<year>2014</year>). <source>Critical Theory and the Digital</source>. Critical Theory and Contemporary Society. <publisher-name>Bloomsbury Academic</publisher-name>.</mixed-citation></ref>
<ref id="B6"><mixed-citation publication-type="webpage"><string-name><surname>Blei</surname>, <given-names>David M.</given-names></string-name>, <string-name><given-names>Andrew Y.</given-names> <surname>Ng</surname></string-name>, and <string-name><given-names>Michael I.</given-names> <surname>Jordan</surname></string-name> (<year>2003</year>). <article-title>&#8220;Latent Dirichlet Allocation&#8221;</article-title>. In: <source>Journal of Machine Learning Research</source> <volume>3</volume> (<issue>1</issue>), <fpage>993</fpage>&#8211;<lpage>1022</lpage>. <uri>https://www.jmlr.org/papers/volume3/blei03a/blei03a.pdf</uri> (visited on 10/03/2024).</mixed-citation></ref>
<ref id="B7"><mixed-citation publication-type="journal"><string-name><surname>Bode</surname>, <given-names>Katherine</given-names></string-name> (<year>2023</year>). <article-title>&#8220;What&#8217;s the Matter with Computational Literary Studies?&#8221;</article-title> In: <source>Critical Inquiry</source> <volume>49</volume> (<issue>4</issue>), <fpage>507</fpage>&#8211;<lpage>529</lpage>. <pub-id pub-id-type="doi">10.1086/724943</pub-id>.</mixed-citation></ref>
<ref id="B8"><mixed-citation publication-type="webpage"><string-name><surname>Boot</surname>, <given-names>Peter</given-names></string-name> (<year>2017</year>). <article-title>&#8220;A Database of Online Book Response and the Nature of the Literary Thriller&#8221;</article-title>. In: <source>Book of Abstracts of DH 2017</source>. <uri>https://dh2017.adho.org/abstracts/208/208.pdf</uri> (visited on 10/07/2024).</mixed-citation></ref>
<ref id="B9"><mixed-citation publication-type="journal"><string-name><surname>Boot</surname>, <given-names>Peter</given-names></string-name> and <string-name><given-names>Marijn</given-names> <surname>Koolen</surname></string-name> (<year>2020</year>). <article-title>&#8220;Captivating, Splendid or Instructive? Assessing the Impact of Reading in Online Book Reviews&#8221;</article-title>. In: <source>Scientific Study of Literature</source> <volume>10</volume> (<issue>1</issue>), <fpage>35</fpage>&#8211;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1075/ssol.20003.boo</pub-id>.</mixed-citation></ref>
<ref id="B10"><mixed-citation publication-type="journal"><string-name><surname>Burrows</surname>, <given-names>John</given-names></string-name> (<year>2006</year>). <article-title>&#8220;All the Way through: Testing for Authorship in Different Frequency Strata&#8221;</article-title>. In: <source>Literary and Linguistic Computing</source> <volume>22</volume> (<issue>1</issue>), <fpage>27</fpage>&#8211;<lpage>47</lpage>.</mixed-citation></ref>
<ref id="B11"><mixed-citation publication-type="journal"><string-name><surname>Cranenburgh</surname>, <given-names>Andreas van</given-names></string-name>, <string-name><given-names>Karina</given-names> <surname>van Dalen-Oskam</surname></string-name>, and <string-name><given-names>Joris van</given-names> <surname>Zundert</surname></string-name> (<year>2019</year>). <article-title>&#8220;Vector Space Explorations of Literary Language&#8221;</article-title>. In: <source>Language Resources and Evaluation</source> <volume>53</volume> (<issue>4</issue>), <fpage>625</fpage>&#8211;<lpage>650</lpage>. <pub-id pub-id-type="doi">10.1007/s10579-018-09442-4</pub-id>.</mixed-citation></ref>
<ref id="B12"><mixed-citation publication-type="book"><string-name><surname>Culpeper</surname>, <given-names>Jonathan</given-names></string-name> and <string-name><given-names>Jane</given-names> <surname>Demmen</surname></string-name> (<year>2015</year>). <chapter-title>&#8220;Keywords&#8221;</chapter-title>. In: <source>The Cambridge Handbook of English Corpus Linguistics</source>. <publisher-name>Cambridge University Press</publisher-name>, <fpage>90</fpage>&#8211;<lpage>105</lpage>.</mixed-citation></ref>
<ref id="B13"><mixed-citation publication-type="webpage"><string-name><surname>Darbyshire</surname>, <given-names>Madison</given-names></string-name> (<year>2023</year>). <article-title>&#8220;Hot Stuff: Why Readers Fell in Love with Romance Novels&#8221;</article-title>. In: <source>Financial Times</source>. <uri>https://www.ft.com/content/0001f781-4927-4780-b46c-3a9f15dffe78</uri> (visited on 10/03/2024).</mixed-citation></ref>
<ref id="B14"><mixed-citation publication-type="webpage"><string-name><surname>Du</surname>, <given-names>Keli</given-names></string-name>, <string-name><given-names>Julia</given-names> <surname>Dudar</surname></string-name>, <string-name><given-names>Cora</given-names> <surname>Rok</surname></string-name>, and <string-name><given-names>Christof</given-names> <surname>Sch&#246;ch</surname></string-name> (<year>2021</year>). <article-title>&#8220;Zeta &amp; Eta: An Exploration and Evaluation of Two Dispersion-based Measures of Distinctiveness&#8221;</article-title>. In: <source>Proceedings of Computational Humanities Research</source>, <fpage>181</fpage>&#8211;<lpage>194</lpage>. <uri>https://ceur-ws.org/Vol-2989/short_paper11.pdf</uri> (visited on 10/03/2024).</mixed-citation></ref>
<ref id="B15"><mixed-citation publication-type="journal"><string-name><surname>Du</surname>, <given-names>Keli</given-names></string-name>, <string-name><given-names>Julia</given-names> <surname>Dudar</surname></string-name>, and <string-name><given-names>Christof</given-names> <surname>Sch&#246;ch</surname></string-name> (<year>2022</year>). <article-title>&#8220;Evaluation of Measures of Distinctiveness. Classification of Literary Texts on the Basis of Distinctive Words&#8221;</article-title>. In: <source>Journal of Computational Literary Studies</source> <volume>1</volume> (<issue>1</issue>). <pub-id pub-id-type="doi">10.48694/jcls.102</pub-id>.</mixed-citation></ref>
<ref id="B16"><mixed-citation publication-type="journal"><string-name><surname>Dunning</surname>, <given-names>Ted</given-names></string-name> (<year>1994</year>). <article-title>&#8220;Accurate Methods for the Statistics of Surprise and Coincidence&#8221;</article-title>. In: <source>Computational Linguistics</source> <volume>19</volume> (<issue>1</issue>), <fpage>61</fpage>&#8211;<lpage>74</lpage>.</mixed-citation></ref>
<ref id="B17"><mixed-citation publication-type="journal"><string-name><surname>Fialho</surname>, <given-names>Olivia</given-names></string-name> (<year>2019</year>). <article-title>&#8220;What Is Literature for? The Role of Transformative Reading&#8221;</article-title>. In: <source>Cogent Arts &amp; Humanities</source> <volume>6</volume> (<issue>1</issue>). Ed. by <string-name><given-names>Anezka</given-names> <surname>Kuzmicova</surname></string-name>. <pub-id pub-id-type="doi">10.1080/23311983.2019.1692532</pub-id>.</mixed-citation></ref>
<ref id="B18"><mixed-citation publication-type="book"><string-name><surname>Gabrielatos</surname>, <given-names>Costas</given-names></string-name> (<year>2018</year>). <chapter-title>&#8220;Keyness Analysis&#8221;</chapter-title>. In: <source>Corpus Approaches to Discourse: A Critical Review</source>. <publisher-name>Routledge</publisher-name>, <fpage>225</fpage>&#8211;<lpage>258</lpage>.</mixed-citation></ref>
<ref id="B19"><mixed-citation publication-type="book"><string-name><surname>Gitelman</surname>, <given-names>Lisa</given-names></string-name> (<year>2013</year>). <source>&#8216;Raw Data&#8217; Is an Oxymoron</source>. <publisher-name>MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.7551/mitpress/9302.001.0001</pub-id>.</mixed-citation></ref>
<ref id="B20"><mixed-citation publication-type="webpage"><string-name><surname>Gries</surname>, <given-names>Stefan Th.</given-names></string-name> (<year>2008</year>). <article-title>&#8220;Dispersions and Adjusted Frequencies in Corpora&#8221;</article-title>. In: <source>International Journal of Corpus Linguistics</source> <volume>13</volume> (<issue>4</issue>), <fpage>403</fpage>&#8211;<lpage>437</lpage>. <uri>https://www.stgries.info/research/2008_STG_Dispersion_IJCL.pdf</uri> (visited on 10/07/2024).</mixed-citation></ref>
<ref id="B21"><mixed-citation publication-type="webpage"><string-name><surname>Hickman</surname>, <given-names>Miranda B.</given-names></string-name> (<year>2012</year>). <chapter-title>&#8220;Introduction: Rereading the New Criticism&#8221;</chapter-title>. In: <source>Rereading the New Criticism</source>. Ed. by <string-name><given-names>Miranda B.</given-names> <surname>Hickman</surname></string-name> and <string-name><given-names>John D.</given-names> <surname>McIntyre</surname></string-name>. <publisher-name>Ohio State University Press</publisher-name>, <fpage>1</fpage>&#8211;<lpage>21</lpage>. <uri>https://core.ac.uk/download/pdf/159569564.pdf</uri> (visited on 10/07/2024).</mixed-citation></ref>
<ref id="B22"><mixed-citation publication-type="webpage"><string-name><surname>Koolen</surname>, <given-names>Marijn</given-names></string-name>, <string-name><given-names>Peter</given-names> <surname>Boot</surname></string-name>, and <string-name><given-names>Joris van</given-names> <surname>Zundert</surname></string-name> (<year>2020</year>). <article-title>&#8220;Online Book Reviews and the Computational Modelling of Reading Impact&#8221;</article-title>. In: <source>Proceedings of the Workshop on Computational Humanities Research</source>, <fpage>149</fpage>&#8211;<lpage>169</lpage>. <uri>http://ceur-ws.org/Vol-2723/long13.pdf</uri> (visited on 10/03/2024).</mixed-citation></ref>
<ref id="B23"><mixed-citation publication-type="webpage"><string-name><surname>Koolen</surname>, <given-names>Marijn</given-names></string-name>, <string-name><given-names>Olivia</given-names> <surname>Fialho</surname></string-name>, <string-name><given-names>Julia</given-names> <surname>Neugarten</surname></string-name>, <string-name><given-names>Joris</given-names> <surname>van Zundert</surname></string-name>, <string-name><given-names>Willem</given-names> <surname>van Hage</surname></string-name>, <string-name><given-names>Ole</given-names> <surname>Mussmann</surname></string-name>, and <string-name><given-names>Peter</given-names> <surname>Boot</surname></string-name> (<year>2023</year>). <article-title>&#8220;How Can Online Book Reviews Validate Empirical In-depth Fiction Reading Typologies?&#8221;</article-title> In: <source>IGEL 2023: Rhythm, Speed, Path: Spatiotemporal Experiences in Narrative, Poetry, and Drama</source>. <uri>https://discourse.igelsociety.org/t/how-can-online-book-reviews-validate-empirical-in-depth-fiction-reading-typologies/370</uri> (visited on 01/16/2024).</mixed-citation></ref>
<ref id="B24"><mixed-citation publication-type="journal"><string-name><surname>Laurino Dos Santos</surname>, <given-names>Henrique</given-names></string-name> and <string-name><given-names>Jonah</given-names> <surname>Berger</surname></string-name> (<year>2022</year>). <article-title>&#8220;The Speed of Stories: Semantic Progression and Narrative Success&#8221;</article-title>. In: <source>Journal of Experimental Psychology. General</source> <volume>151</volume> (<issue>8</issue>), <fpage>1833</fpage>&#8211;<lpage>1842</lpage>. <pub-id pub-id-type="doi">10.1037/xge0001171</pub-id>.</mixed-citation></ref>
<ref id="B25"><mixed-citation publication-type="journal"><string-name><surname>Lijffijt</surname>, <given-names>Jefrey</given-names></string-name>, <string-name><given-names>Terttu</given-names> <surname>Nevalainen</surname></string-name>, <string-name><given-names>Tanja</given-names> <surname>S&#228;ily</surname></string-name>, <string-name><given-names>Panagiotis</given-names> <surname>Papapetrou</surname></string-name>, <string-name><given-names>Kai</given-names> <surname>Puolam&#228;ki</surname></string-name>, and <string-name><given-names>Heikki</given-names> <surname>Mannila</surname></string-name> (<year>2016</year>). <article-title>&#8220;Significance Testing of Word Frequencies in Corpora&#8221;</article-title>. In: <source>Digital Scholarship in the Humanities</source> <volume>31</volume> (<issue>2</issue>), <fpage>374</fpage>&#8211;<lpage>397</lpage>. <pub-id pub-id-type="doi">10.1093/llc/fqu064</pub-id>.</mixed-citation></ref>
<ref id="B26"><mixed-citation publication-type="journal"><string-name><surname>Loi</surname>, <given-names>Christina</given-names></string-name>, <string-name><given-names>Frank</given-names> <surname>Hakemulder</surname></string-name>, <string-name><given-names>Moniek</given-names> <surname>Kuijpers</surname></string-name>, and <string-name><given-names>Gerhard</given-names> <surname>Lauer</surname></string-name> (<year>2023</year>). <article-title>&#8220;On How Fiction Impacts the Self-concept: Transformative Reading Experiences and Storyworld Possible Selves&#8221;</article-title>. In: <source>Scientific Study of Literature</source> <volume>12</volume> (<issue>1</issue>), <fpage>44</fpage>&#8211;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.61645/ssol.181</pub-id>.</mixed-citation></ref>
<ref id="B27"><mixed-citation publication-type="book"><string-name><surname>Manovich</surname>, <given-names>Lev</given-names></string-name> (<year>2013</year>). <source>Software Takes Command</source>. International Texts in Critical Media Aesthestics. <publisher-name>Bloomsbury Academic</publisher-name>.</mixed-citation></ref>
<ref id="B28"><mixed-citation publication-type="journal"><string-name><surname>Miall</surname>, <given-names>David S.</given-names></string-name> and <string-name><given-names>Don</given-names> <surname>Kuiken</surname></string-name> (<year>1994</year>). <article-title>&#8220;Beyond Text Theory: Understanding Literary Response&#8221;</article-title>. In: <source>Discourse Processes</source> <volume>17</volume> (<issue>3</issue>), <fpage>337</fpage>&#8211;<lpage>352</lpage>. <pub-id pub-id-type="doi">10.1080/01638539409544873</pub-id>.</mixed-citation></ref>
<ref id="B29"><mixed-citation publication-type="webpage"><string-name><surname>Moreira</surname>, <given-names>Pascale</given-names></string-name>, <string-name><given-names>Yuri</given-names> <surname>Bizzoni</surname></string-name>, <string-name><given-names>Kristoffer</given-names> <surname>Nielbo</surname></string-name>, <string-name><given-names>Ida Marie</given-names> <surname>Lassen</surname></string-name>, and <string-name><given-names>Mads</given-names> <surname>Thomsen</surname></string-name> (<year>2023</year>). <article-title>&#8220;Modeling Readers&#8217; Appreciation of Literary Narratives Through Sentiment Arcs and Semantic Profiles&#8221;</article-title>. In: <source>Proceedings of the the 5th Workshop on Narrative Understanding</source>, <fpage>25</fpage>&#8211;<lpage>35</lpage>. <uri>https://aclanthology.org/2023.wnu-1.5</uri> (visited on 01/22/2024).</mixed-citation></ref>
<ref id="B30"><mixed-citation publication-type="journal"><string-name><surname>Nguyen</surname>, <given-names>Minh Van</given-names></string-name>, <string-name><given-names>Viet</given-names> <surname>Lai</surname></string-name>, <string-name><given-names>Amir</given-names> <surname>Pouran Ben Veyseh</surname></string-name>, and <string-name><given-names>Thien Huu</given-names> <surname>Nguyen</surname></string-name> (<year>2021</year>). <article-title>&#8220;Trankit: A Light-Weight Transformer-based Toolkit for Multilingual Natural Language Processing&#8221;</article-title>. In: <source>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations</source>. <pub-id pub-id-type="doi">10.18653/v1/2021.eacl-demos.10</pub-id>.</mixed-citation></ref>
<ref id="B31"><mixed-citation publication-type="book"><string-name><surname>Paquot</surname>, <given-names>Magali</given-names></string-name> and <string-name><given-names>Yves</given-names> <surname>Bestgen</surname></string-name> (<year>2009</year>). <chapter-title>&#8220;Distinctive Words in Academic Writing: A Comparison of Three Statistical Tests for Keyword Extraction&#8221;</chapter-title>. In: <source>Corpora: Pragmatics and Discourse</source>. <publisher-name>Brill</publisher-name>, <fpage>247</fpage>&#8211;<lpage>269</lpage>. <pub-id pub-id-type="doi">10.1163/9789042029101_014</pub-id>.</mixed-citation></ref>
<ref id="B32"><mixed-citation publication-type="journal"><string-name><surname>Pichler</surname>, <given-names>Axel</given-names></string-name> and <string-name><given-names>Nils</given-names> <surname>Reiter</surname></string-name> (<year>2022</year>). <article-title>&#8220;From Concepts to Texts and Back: Operationalization as a Core Activity of Digital Humanities&#8221;</article-title>. In: <source>Journal of Cultural Analytics</source> <volume>7</volume> (<issue>4</issue>). <pub-id pub-id-type="doi">10.22148/001c.57195</pub-id>.</mixed-citation></ref>
<ref id="B33"><mixed-citation publication-type="webpage"><string-name><surname>Prescott</surname>, <given-names>Andrew</given-names></string-name> (<year>2023</year>). <article-title>&#8220;Bias in Big Data, Machine Learning and AI: What Lessons for the Digital Humanities?&#8221;</article-title> In: <source>Digital Humanities Quarterly</source> <volume>17</volume> (<issue>2</issue>). <uri>https://www.digitalhumanities.org/dhq/vol/17/2/000689/000689.html</uri> (visited on 01/16/2024).</mixed-citation></ref>
<ref id="B34"><mixed-citation publication-type="webpage"><string-name><surname>Rawson</surname>, <given-names>Katie</given-names></string-name> and <string-name><given-names>Trevor</given-names> <surname>Mu&#241;oz</surname></string-name> (<year>2016</year>). <source>Against Cleaning</source>. Project Blog. <uri>http://curatingmenus.org/articles/against-cleaning/</uri> (visited on 09/30/2016).</mixed-citation></ref>
<ref id="B35"><mixed-citation publication-type="book"><string-name><surname>Regis</surname>, <given-names>Pamela</given-names></string-name> (<year>2003</year>). <source>A Natural History of the Romance Novel</source>. <publisher-name>University of Pennsylvania Press</publisher-name>.</mixed-citation></ref>
<ref id="B36"><mixed-citation publication-type="webpage"><string-name><surname>Sch&#246;ch</surname>, <given-names>Christof</given-names></string-name> (<year>2017</year>). <article-title>&#8220;Topic Modeling Genre: An Exploration of French Classical and Enlightenment Drama&#8221;</article-title>. In: <source>Digital Humanities Quarterly</source> <volume>11</volume> (<issue>2</issue>). <uri>https://www.digitalhumanities.org/dhq/vol/11/2/000291/000291.html</uri> (visited on 10/07/2024).</mixed-citation></ref>
<ref id="B37"><mixed-citation publication-type="journal"><string-name><surname>Sobchuk</surname>, <given-names>Oleg</given-names></string-name> and <string-name><given-names>Artjoms</given-names> <surname>&#352;e&#316;a</surname></string-name> (<year>2023</year>). <article-title>&#8220;Computational Thematics: Comparing Algorithms for Clustering the Genres of Literary Fiction&#8221;</article-title>. In: <source>arXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2305.11251</pub-id>.</mixed-citation></ref>
<ref id="B38"><mixed-citation publication-type="webpage"><string-name><surname>Thompson</surname>, <given-names>Laure</given-names></string-name> and <string-name><given-names>David</given-names> <surname>Mimno</surname></string-name> (<year>2018</year>). <article-title>&#8220;Authorless Topic Models: Biasing Models Away from Known Structure&#8221;</article-title>. In: <source>Proceedings of the 27th International Conference on Computational Linguistics</source>, <fpage>3903</fpage>&#8211;<lpage>3914</lpage>. <uri>https://aclanthology.org/C18-1329</uri> (visited on 10/03/2024).</mixed-citation></ref>
<ref id="B39"><mixed-citation publication-type="journal"><string-name><surname>Toubia</surname>, <given-names>Olivier</given-names></string-name>, <string-name><given-names>Jonah</given-names> <surname>Berger</surname></string-name>, and <string-name><given-names>Jehoshua</given-names> <surname>Eliashberg</surname></string-name> (<year>2021</year>). <article-title>&#8220;How Quantifying the Shape of Stories Predicts Their Success&#8221;</article-title>. In: <source>Proceedings of the National Academy of Sciences</source> <volume>118</volume> (<issue>26</issue>), <fpage>1</fpage>&#8211;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.2011695118</pub-id>.</mixed-citation></ref>
<ref id="B40"><mixed-citation publication-type="webpage"><string-name><surname>Uglanova</surname>, <given-names>Inna</given-names></string-name> and <string-name><given-names>Evelyn</given-names> <surname>Gius</surname></string-name> (<year>2020</year>). <article-title>&#8220;The Order of Things. A Study on Topic Modelling of Literary Texts&#8221;</article-title>. In: <source>Proceedings of the Workshop on Computational Humanities Research</source>, <fpage>57</fpage>&#8211;<lpage>76</lpage>. <uri>https://ceur-ws.org/Vol-2723/long7.pdf</uri> (visited on 10/03/2024).</mixed-citation></ref>
<ref id="B41"><mixed-citation publication-type="webpage"><string-name><surname>Warnock</surname>, <given-names>John</given-names></string-name> (<year>1978</year>). <article-title>&#8220;A Theory of Discourse, by James L. Kinneavy. (Review)&#8221;</article-title>. In: <source>Style</source> <volume>12</volume> (<issue>1</issue>), <fpage>52</fpage>&#8211;<lpage>54</lpage>. <uri>https://www.jstor.org/stable/45109026</uri> (visited on 01/16/2024).</mixed-citation></ref>
<ref id="B42"><mixed-citation publication-type="webpage"><string-name><surname>Wimsatt</surname>, <given-names>William K.</given-names></string-name> (<year>1954</year>). <chapter-title>&#8220;The Intentional Fallacy&#8221;</chapter-title>. In: <source>The Verbal Icon: Studies in the Meaning of Poetry</source>. <publisher-name>University Press of Kentucky</publisher-name>, <fpage>3</fpage>&#8211;<lpage>20</lpage>. <uri>https://www.sas.upenn.edu/cavitch/pdf-library/WimsattBeardsley_Intentional.pdf</uri> (visited on 10/07/2024).</mixed-citation></ref>
<ref id="B43"><mixed-citation publication-type="webpage"><string-name><surname>Zundert</surname>, <given-names>Joris van</given-names></string-name>, <string-name><given-names>Marijn</given-names> <surname>Koolen</surname></string-name>, and <string-name><given-names>Karina van</given-names> <surname>Dalen-Oskam</surname></string-name> (<year>2018</year>). <article-title>&#8220;Predicting Prose that Sells: Issues of Open Data in a Case of Applied Machine Learning&#8221;</article-title>. In: <source>JADH 2018 &#8216;Leveraging Open Data&#8217;: Proceedings of the 8th Conference of Japanese Association for Digital Humanities</source>, <fpage>175</fpage>&#8211;<lpage>177</lpage>. <uri>https://conf2018.jadh.org/files/Proceedings_JADH2018_rev0911.pdf</uri> (visited on 11/07/2018).</mixed-citation></ref>
<ref id="B44"><mixed-citation publication-type="webpage"><string-name><surname>Zundert</surname>, <given-names>Joris van</given-names></string-name>, <string-name><given-names>Marijn</given-names> <surname>Koolen</surname></string-name>, <string-name><given-names>Julia</given-names> <surname>Neugarten</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Boot</surname></string-name>, <string-name><given-names>Willem</given-names> <surname>van Hage</surname></string-name>, and <string-name><given-names>Ole</given-names> <surname>Mussmann</surname></string-name> (<year>2022</year>). <article-title>&#8220;What Do We Talk About When We Talk About Topic?&#8221;</article-title> In: <source>Proceedings of Computational Humanities Research</source>, <fpage>398</fpage>&#8211;<lpage>410</lpage>. <uri>https://ceur-ws.org/Vol-3290/</uri> (visited on 11/22/2023).</mixed-citation></ref>
</ref-list>
<sec id="A1">
<title>A. Mapping NUR Codes to Genre Labels</title>
<p>The complete mapping of NUR codes to genre labels is shown in <xref ref-type="table" rid="AT2">Table 2</xref>.</p>
<table-wrap id="AT2">
<caption>
<p><bold>Table 2:</bold> The selected NUR codes of novels in our dataset of 18,885 novels and their mapping to genres.</p>
</caption>
<table>
<thead>
<tr>
<td align="right" valign="top">NUR code</td>
<td align="left" valign="top">NUR label</td>
<td align="left" valign="top">Genre label</td>
</tr>
</thead>
<tbody>
<tr>
<td align="right" valign="top">280</td>
<td align="left" valign="top">Children&#8217;s Fiction general</td>
<td align="left" valign="top">Children&#8217;s fiction</td>
</tr>
<tr>
<td align="right" valign="top">281</td>
<td align="left" valign="top">Children&#8217;s fiction 4&#8211;6 years</td>
<td align="left" valign="top">Children&#8217;s fiction</td>
</tr>
<tr>
<td align="right" valign="top">282</td>
<td align="left" valign="top">Children&#8217;s fiction 7&#8211;9 years</td>
<td align="left" valign="top">Children&#8217;s fiction</td>
</tr>
<tr>
<td align="right" valign="top">283</td>
<td align="left" valign="top">Children&#8217;s fiction 10&#8211;12 years</td>
<td align="left" valign="top">Children&#8217;s fiction</td>
</tr>
<tr>
<td align="right" valign="top">284</td>
<td align="left" valign="top">Children&#8217;s fiction 13&#8211;15 years</td>
<td align="left" valign="top">Young adult</td>
</tr>
<tr>
<td align="right" valign="top">285</td>
<td align="left" valign="top">Children&#8217;s fiction 15+</td>
<td align="left" valign="top">Young adult</td>
</tr>
<tr>
<td align="right" valign="top">300</td>
<td align="left" valign="top">Literary fiction general</td>
<td align="left" valign="top">Literary fiction</td>
</tr>
<tr>
<td align="right" valign="top">301</td>
<td align="left" valign="top">Literary fiction Dutch</td>
<td align="left" valign="top">Literary fiction</td>
</tr>
<tr>
<td align="right" valign="top">302</td>
<td align="left" valign="top">Literary fiction translated</td>
<td align="left" valign="top">Literary fiction</td>
</tr>
<tr>
<td align="right" valign="top">305</td>
<td align="left" valign="top">Literary thriller</td>
<td align="left" valign="top">Literary thriller</td>
</tr>
<tr>
<td align="right" valign="top">312</td>
<td align="left" valign="top">Pockets popular fiction</td>
<td align="left" valign="top">Literary fiction</td>
</tr>
<tr>
<td align="right" valign="top">313</td>
<td align="left" valign="top">Pockets suspense</td>
<td align="left" valign="top">Suspense</td>
</tr>
<tr>
<td align="right" valign="top">330</td>
<td align="left" valign="top">Suspense general</td>
<td align="left" valign="top">Suspense</td>
</tr>
<tr>
<td align="right" valign="top">331</td>
<td align="left" valign="top">Detective</td>
<td align="left" valign="top">Suspense</td>
</tr>
<tr>
<td align="right" valign="top">332</td>
<td align="left" valign="top">Thriller</td>
<td align="left" valign="top">Suspense</td>
</tr>
<tr>
<td align="right" valign="top">334</td>
<td align="left" valign="top">Fantasy</td>
<td align="left" valign="top">Fantasy fiction</td>
</tr>
<tr>
<td align="right" valign="top">339</td>
<td align="left" valign="top">True crime</td>
<td align="left" valign="top">Suspense</td>
</tr>
<tr>
<td align="right" valign="top">342</td>
<td align="left" valign="top">Historical novel (popular)</td>
<td align="left" valign="top">Historical fiction</td>
</tr>
<tr>
<td align="right" valign="top">343</td>
<td align="left" valign="top">Romance</td>
<td align="left" valign="top">Romance</td>
</tr>
<tr>
<td align="right" valign="top">344</td>
<td align="left" valign="top">Regional and family novel</td>
<td align="left" valign="top">Regional fiction</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="A2">
<title>B. Overlap between Themes in Terms of Shared Books</title>
<p>The topic modeling process assigns each book to a single topic, but because individual topics can be linked to multiple themes, their books are also linked to multiple themes. As a consequence, themes share books and reviews, and some pairs of themes may have larger overlap than others. This overlap between themes is shown for pairs of themes where for one theme at least 25% of the books for one theme are shared by the other theme.</p>
<table-wrap id="AT3">
<caption>
<p><bold>Table 3:</bold> Overlap in books between themes, for themes where one theme shares at least 25% of the books with the other theme.</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"></td>
<td align="right" valign="top"></td>
<td align="left" valign="top"></td>
<td align="right" valign="top"></td>
<td align="right" valign="bottom">Book</td>
<td align="center" valign="bottom" colspan="2">Books</td>
</tr>
<tr>
<td align="left" valign="top">Theme 1</td>
<td align="right" valign="top">Share 1</td>
<td align="left" valign="top">Theme 2</td>
<td align="right" valign="top">Share 2</td>
<td align="right" valign="bottom">Overlap</td>
<td align="right" valign="top">Theme 1</td>
<td align="right" valign="top">Theme 2</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">crime</td>
<td align="right" valign="top">0.33</td>
<td align="left" valign="top">geo. &amp; setting</td>
<td align="right" valign="top">0.14</td>
<td align="right" valign="top">619</td>
<td align="right" valign="top">1,899</td>
<td align="right" valign="top">4,317</td>
</tr>
<tr>
<td align="left" valign="top">culture</td>
<td align="right" valign="top">0.49</td>
<td align="left" valign="top">geo. &amp; setting</td>
<td align="right" valign="top">0.40</td>
<td align="right" valign="top">1,713</td>
<td align="right" valign="top">3,524</td>
<td align="right" valign="top">4,317</td>
</tr>
<tr>
<td align="left" valign="top">econ. &amp; work</td>
<td align="right" valign="top">0.36</td>
<td align="left" valign="top">behav./feelings</td>
<td align="right" valign="top">0.12</td>
<td align="right" valign="top">446</td>
<td align="right" valign="top">1,232</td>
<td align="right" valign="top">3,860</td>
</tr>
<tr>
<td align="left" valign="top">econ. &amp; work</td>
<td align="right" valign="top">0.30</td>
<td align="left" valign="top">society</td>
<td align="right" valign="top">0.44</td>
<td align="right" valign="top">371</td>
<td align="right" valign="top">1,232</td>
<td align="right" valign="top">851</td>
</tr>
<tr>
<td align="left" valign="top">econ. &amp; work</td>
<td align="right" valign="top">0.25</td>
<td align="left" valign="top">politics</td>
<td align="right" valign="top">0.49</td>
<td align="right" valign="top">310</td>
<td align="right" valign="top">1,232</td>
<td align="right" valign="top">634</td>
</tr>
<tr>
<td align="left" valign="top">family</td>
<td align="right" valign="top">0.65</td>
<td align="left" valign="top">behav./feelings</td>
<td align="right" valign="top">0.08</td>
<td align="right" valign="top">324</td>
<td align="right" valign="top">498</td>
<td align="right" valign="top">3,860</td>
</tr>
<tr>
<td align="left" valign="top">family</td>
<td align="right" valign="top">0.30</td>
<td align="left" valign="top">culture</td>
<td align="right" valign="top">0.04</td>
<td align="right" valign="top">151</td>
<td align="right" valign="top">498</td>
<td align="right" valign="top">3,524</td>
</tr>
<tr>
<td align="left" valign="top">geo. &amp; setting</td>
<td align="right" valign="top">0.40</td>
<td align="left" valign="top">culture</td>
<td align="right" valign="top">0.49</td>
<td align="right" valign="top">1,713</td>
<td align="right" valign="top">4,317</td>
<td align="right" valign="top">3,524</td>
</tr>
<tr>
<td align="left" valign="top">history</td>
<td align="right" valign="top">0.51</td>
<td align="left" valign="top">geo. &amp; setting</td>
<td align="right" valign="top">0.24</td>
<td align="right" valign="top">1,038</td>
<td align="right" valign="top">2,020</td>
<td align="right" valign="top">4,317</td>
</tr>
<tr>
<td align="left" valign="top">history</td>
<td align="right" valign="top">0.31</td>
<td align="left" valign="top">war</td>
<td align="right" valign="top">0.65</td>
<td align="right" valign="top">622</td>
<td align="right" valign="top">2,020</td>
<td align="right" valign="top">952</td>
</tr>
<tr>
<td align="left" valign="top">life st. &amp; sport</td>
<td align="right" valign="top">0.31</td>
<td align="left" valign="top">medi./health</td>
<td align="right" valign="top">0.20</td>
<td align="right" valign="top">216</td>
<td align="right" valign="top">702</td>
<td align="right" valign="top">1,058</td>
</tr>
<tr>
<td align="left" valign="top">politics</td>
<td align="right" valign="top">0.49</td>
<td align="left" valign="top">econ. &amp; work</td>
<td align="right" valign="top">0.25</td>
<td align="right" valign="top">310</td>
<td align="right" valign="top">634</td>
<td align="right" valign="top">1,232</td>
</tr>
<tr>
<td align="left" valign="top">politics</td>
<td align="right" valign="top">0.49</td>
<td align="left" valign="top">society</td>
<td align="right" valign="top">0.36</td>
<td align="right" valign="top">310</td>
<td align="right" valign="top">634</td>
<td align="right" valign="top">851</td>
</tr>
<tr>
<td align="left" valign="top">society</td>
<td align="right" valign="top">0.44</td>
<td align="left" valign="top">econ. &amp; work</td>
<td align="right" valign="top">0.30</td>
<td align="right" valign="top">371</td>
<td align="right" valign="top">851</td>
<td align="right" valign="top">1,232</td>
</tr>
<tr>
<td align="left" valign="top">society</td>
<td align="right" valign="top">0.36</td>
<td align="left" valign="top">politics</td>
<td align="right" valign="top">0.49</td>
<td align="right" valign="top">310</td>
<td align="right" valign="top">851</td>
<td align="right" valign="top">634</td>
</tr>
<tr>
<td align="left" valign="top">war</td>
<td align="right" valign="top">0.65</td>
<td align="left" valign="top">history</td>
<td align="right" valign="top">0.31</td>
<td align="right" valign="top">622</td>
<td align="right" valign="top">952</td>
<td align="right" valign="top">2,020</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="A3">
<title>C. Correlations between Themes in Terms of Impact</title>
<p>The correlations between themes in terms of the percent difference (%Diff) per impact term for generic <italic>Affect, Narrative</italic>, and <italic>Aesthetics</italic> are shown in <xref ref-type="fig" rid="F13">Figure 13</xref>, <xref ref-type="fig" rid="F14">Figure 14</xref>, and <xref ref-type="fig" rid="F15">Figure 15</xref>, respectively.</p>
<fig id="F13">
<caption>
<p><bold>Figure 13:</bold> Percent different correlations between themes based on general <italic>Affect</italic> terms.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g13.png"/>
</fig>
<fig id="F14">
<caption>
<p><bold>Figure 14:</bold> Percent different correlations between themes based on <italic>Narrative</italic> impact terms.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g14.png"/>
</fig>
<fig id="F15">
<caption>
<p><bold>Figure 15:</bold> Percent different correlations between themes based on general <italic>Aesthetic</italic> impact terms.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="jcls-3-1-3927-g15.png"/>
</fig>
</sec>
</back>
</article>