“Books are well written, or badly written.”
- Oscar Wilde
“There is nothing either good or bad, but thinking makes it so.”
- Hamlet (William Shakespeare)
1. Introduction
How much of a novel’s merit is debatable? To what degree is good fiction in the eye of the beholder? Newspaper articles and awards committees often provide strong opinions, thereby steering the economic fate of authors and determining what the public considers to be good literature (Chuey et al. 2024). In this paper, we set out to empirically disentangle the relative contributions of book and reader characteristics in determining online book reviews.
2. Subjectivity in Book Ratings
Whether discussing architecture, music, or written works, people have always debated whether beauty comes down to personal taste or is, at least sometimes, undeniable (Ball 2004; Gessey-Jones et al. 2020; Hume [1757] 2017; Hyman 2002; Pelowski et al. 2017).
In the world of fiction writing, most creators, editors, and publishing professionals agree that different books appeal to different readers (Alharthi et al. 2018; Baumard et al. 2022; Waples 1931). As a clear sign of this, people’s literary taste changes over their lifespan; the very same reader might love a book as a child but then loathe it a few years later (Reuter 2007). Some books might entertain millions of adoring fans, while being snubbed by professional critics (Bizzoni et al. 2023). Once celebrated books might lose their appeal to modern audiences (Shaffi 2023). All this suggests that a reader’s perspective matters when assessing book quality. But how much?
Bizzoni and colleagues discuss whether mild perspectivism can provide an adequate description of reading enjoyment, where books do vary in their average enjoyability, but differentiable reader groups have unique criteria for books. As examples of reader groups, the authors list professional critics, lay readers, and demographic groups, which vary in their reading appreciation (Bizzoni et al. 2022). In fact, readers themselves report enjoying different literary genres (Herald and Stavole-Carter 2019) and such self-ascribed preferences affect how they search for new books (Mikkonen and Vakkari 2017), which reading communities they join (Kraaykamp and Dijkstra 1999; Thelwall and Kousha 2017), and which recommendations they receive by librarians and literature websites (Benkoussas et al. 2014; Thelwall and Kousha 2017).
The biggest divide in reading enjoyment lies between people claiming that they hate books in general and people who make sweeping declarations to love books (Greaney and Hegarty 1987). Long-term reading enthusiasts often connect books to various positive experiences throughout their life from bonding with parents to escaping their sorrows during difficult times (Carlsen and Sherrill 1988). They appreciate books as stimulants of their imagination, and often form habits around reading (Strommen and Mates 2004). Conversely, book skeptics tend to reject books for reasons of anticipated boredom or exhaustion (Gordon and Lu 2021), or because they genuinely do not think they have the skills to read a book (Unrau et al. 2018). Teachers complain that many students lack patience for books or have been poached by fast-paced, visual media (Poletti et al. 2016). Non-surprisingly, a person’s overall attitude towards books, as well as their self-perception as a reader, primes their subsequent enjoyment of books (Adelson et al. 2019; Ramirez et al. 2019). In sum, there is widespread agreement that a book’s perceived quality depends on characteristics of the reader, at least to some degree.
3. Objectivity in Book Ratings
While the reader’s perspective matters, most people also believe that some texts are truly better than others (Croxson et al. 2021). In fact, books are regularly bestowed with absolutist scores, reviews, and awards that celebrate the ‘best’ literature (National Book Foundation 2024). There are instructional handbooks on how to write better (White and Strunk 2023), lists of common writing mistakes (Mittelmark and Newman 2009), and detailed how-to-guides by famous authors (King 2000).
Most people would likely agree that a book filled with (unintentional) grammatical errors, logical inconsistencies, and written in a challenging font has some avoidable shortcomings. People are so certain that these things are mistakes that students are taught to avoid them in their writing. Admittedly, professional authors can eradicate such superficial errors quite easily; however, the mere observation that one can get better at writing suggests that authors might still differ on the latent dimension of ‘writing skill’. For instance, other common reader complaints, like unlikable or flat characters, are much more abstract and challenging to avoid, and thus more difficult to master than grammatical rule following (Balée 2006; Walsh and Antoniak 2021). An analysis by Kaufman and Kaufman suggests that authors often need more than ten years of professional experience to produce their best work (S. B. Kaufman and J. C. Kaufman 2007).
If authors can become better at writing, this entails that books can differ in their quality, provided that an adherence to writing conventions translates to increased appreciation by the readers. Tankard and Hendrickson tested whether one of the most famous writing tips (Show, do not tell!) actually has a positive influence on reading enjoyment (Tankard and Hendrickson 1996). They found that the descriptive show-versions of sentences (e.g., “Suddenly I awoke in a drenching sweat, my heart racing”) were indeed perceived as more interesting and engaging than the abstract tell-versions (“Suddenly I awoke, frightened”). However, not all writing conventions correlate with book popularity as Boyd and colleagues found no benefit of choosing any narrative structure over another in a sample of 60,000 works of fiction (Boyd et al. 2020).
In sum, it is a common belief, backed by some empirical evidence, that a story can be written well or poorly, meaning that some books might truly be better than others. We are left with two questions. First, to what degree is a book’s perceived quality due to the book, versus the reader? And second, how could one measure a book’s quality?
4. How to Measure Book Quality
It turns out that the aforementioned questions are two sides of the same coin. Both the nature and the measurement of book quality pose the problem of which reviewer to believe. One field that battles this problem is computational literary studies, where machine learning algorithms are employed to predict a book’s success. (Notice the cliched ‘coin’ metaphor in the paragraph’s first sentence. Did it hamper your reading enjoyment, as it is unoriginal, or did you find it useful, as it is succinct?)
While some researchers focus on predicting book sales or download counts (Ganjigunte Ashok et al. 2013; Otten et al. 2019; Silva et al. 2024; Sujo et al. 2021; Wang et al. 2019; Yucesoy et al. 2018), others explicitly forecast canonicity, award wins, or review scores (Barré et al. 2023; Maharjan et al. 2017; Pascake Moreira et al. 2023; Pascale Moreira and Bizzoni 2023). Such prediction models have the power to steer publishers’ investments and therefore determine what the public reads and which authors become successful. However, can we assume that the chosen prediction targets – specifically reader reviews and committee scores – are reliable measures of a book’s quality?
Bizzoni et al. (2022) describe two opposing views in this regard: strong perspectivism and weak perspectivism. The former denotes that even a single review score constitutes an accurate measurement of a book’s quality, but might not generalize to other readers. Thus, a book’s quality should only be measured and predicted for the individual reader, and never in absolute terms. Weak perspectivism, on the other hand, argues that book quality can be measured in absolute terms, but might require averaging across many reviews by different people. It postulates that individual review scores serve as noisy indicators of a book’s latent quality. However, how much idiosyncratic noise overlays the latent book quality remains unknown.
Researchers predicting aesthetics ratings in other fields have quantified the amount of inter-individual variation to distinguish it from other forms of statistical ‘noise’. For instance, Hönekopp (2006) found that judgments of facial beauty are about equally determined by face-characteristics and rater-tendencies. Hehman et al. (2017) extended this finding by showing that some face ratings do depend on the observed face (e.g., perceived happiness) whereas other traits are almost entirely due to the beholder (e.g., perceived creativity). Human-generated art like paintings, architecture, and textiles elicit more person-specific aesthetics ratings than faces and nature scenes (Vessel et al. 2018), a trend that is more pronounced for abstract art than representational art (Leder et al. 2016; Schepman et al. 2015).
When it comes to text evaluations, there is evidence for both subjectivity and objectivity. The quality of news domains, for instance, appears to be rated with high agreement among experts (Lin et al. 2023), whereas reviewers of scientific manuscripts often disagree in their evaluations (Siler et al. 2015).
The appreciation of textual humor appears to be largely subjective (Rosenbusch et al. 2022) and, as textual humor is a form of creative writing, we decided to pre-register our numerical hypotheses – about the relative importance of books and readers – in direct accordance with the confidence intervals for written jokes (cf. Rosenbusch et al. 2022; study 2):
Relative-importance-hypothesis
Differences between raters will account for more than three times as much rating variance as differences between books.
Book-importance-hypothesis
Differences between books will account for five to nine percent of all rating variance.
Rater-importance-hypothesis
Differences between raters will account for thirty to thirty-eight percent of all rating variance.
Next to broadening our knowledge of subjectivity in text evaluations, the presented analyses will provide answers to practical questions like: Should one use book reviews for purchasing decisions? How many reviewers are needed to assess a book’s quality and make fair comparisons between books? And which pieces of information are needed to predict a book’s reception by the public?
5. Method and Materials
Pre-registration, code, and data can be found in section 8 and section 9. We analyze book reviews from one of the largest literary websites: goodreads.com (cf., Walsh and Antoniak 2021). Goodreads users rate books by clicking on a star scale, reaching from one to five stars, or by writing a free text review.
To collect the data, a Python script iterated through random numbers (numpy.random.randint; seeded) and inserted them into the website’s url, stopping after 300,000 numbers. It is unknown how Goodreads assigns URLs and we therefore decided that uniform sampling posed the lowest risk of sampling bias. Given that webscraping is generally disfavored/blocked by website hosts, and given that Goodreads has shut down its API for programatically retrieving book reviews, it is possible that our data collection might not be reproducible in the future. In this case, researchers will be able to reanalyze our numerical, anonymous data. Private accounts, and reviews that were not posted publicly for the entire reading community, were not (and cannot be) accessed. From each random user, up to twenty random reviews were collected. Users with zero or one review were also discarded. No identifying data were collected. The final dataset consists of 566,121 ratings, from 67,012 participants, and across 49,674 books. Users annotated the book corpus with 875 unique genres ranging from very general (‘fiction’) to very specific tags (‘World War II’). In line with our preregistration, no genre subsampling or weighting was performed.
During the analyses, we disentangle two sources of variance, the book and the reviewer, through the use of random effects in multilevel models and the ICC value associated with each source of variance. Thus, we operationalize differences between books and differences between readers as inter-book and inter-reader variance in book ratings, respectively. We further show probabilities of book ratings conditional on other ratings to assess rating consistency across readers and books. Confidence intervals are computed by bootstrap sampling each analysis with ten repetitions and noting the highest and lowest values. Note that we pre-registered 100 rather than ten repetitions, which were unneeded given the large sample size.
6. Results
A multilevel model with random intercepts per book and reader (Bates et al. 2015, 2024) revealed that differences between books account for 3.53% [3.42–3.68] of the rating variance on Goodreads, whereas differences between raters accounted for 30.28% [30.06–30.51]. Reader differences were 8.58 times [8.2–8.87] more impactful than differences between books. Figure 1 illustrates the relative benefit of knowing the rater vs. knowing the book when trying to predict rating scores. The relatively high consistency of ratings from a given reader, compared to the lack of consistency for ratings of the same book, is also highlighted by an extended simulation where we sampled an increasing number of raters and averaged their score to approximate either a single rater’s score, or the book’s average score in our complete dataset (cf. Figure 2).
Figure 1: Top left: conditional probability of a book rating, given another reader’s rating of the same book (If Rater A gave a book one star, there is a 50% chance that Rater B will give it five stars). Bottom left: conditional probability of a book rating, given the same reader’s rating of another book. Top right: consistency between average book ratings (Ns=3), all made about the same book. Bottom right: consistency between average book ratings (Ns=3), all made by the same reader. Takeaway: Two raters rarely agree (tops), but individual raters have consistent tendencies (bottoms).
Figure 2: Left: Pearson correlation of sample ratings with the rating of a single rater. Right: correlation of sample ratings with books’ overall ratings in our dataset. Notice that these correlation coefficients lie far below conventions for reliable quality assessments (Montgomery et al. 2002), despite being inflated by the rating samples forming part of the ‘All Ratings’ average.
6.1 Subsample Analyses
When restricting the book sample to written fiction books (i.e., excluding non-fiction, audiobooks, and image content, see supplementary file genre_filters.csv for all 20 tags that were excluded; remaining N = 26,160), the variances accounted for by book differences (3.23% [3.07-3.39]) and rater differences (29.61% [29.26-29.92]) stayed virtually the same. Note, that this also excluded the 800 edge cases where books had both a fiction and a non-fiction label (e.g., dramatic retellings of war stories). Another sensitivity analysis on the 50% of books with the least amount of ratings, showed a similar impact of books (3.23%; [2.52-4.03]) and an even higher impact of the rater (40.76%; [40.35-41.14]). When removing all users that gave the same rating to all their reviewed books (which includes users with few ratings, user that seemingly use the platform to collect favorite books, and potentially undiscerning book enthusiasts), book differences accounted for (4.3%; [4.06-4.47]), while reader tendencies by definition diminished in importance but were still considerably higher (23.14%; [22.76-23.49]; Nratings = 497,413).
It is possible that the amount of rater agreement is higher within genres, as ‘genre tourists’ might introduce reader-level variance through a less practiced eye for genre-adequate writing. Thus, we restricted the sample to book ratings from readers who had rated books from the same genre positively (4 or 5 stars) at least twice. We operationalized ‘same genre’ books as having at least three shared tags out of the twenty most common genre tags, or the exact same set of genre tags (in case the book had less than three tags). The resulting dataset of 433,149 ‘within-genre-ratings’ showed that book differences still accounted for fairly little variance (3.86%; [3.68-4.09]), while the importance of rater tendencies diminished (22.63%; [22.37-22.97]), potentially due to overall higher ratings. Lastly, it is conceivable that readers who are practiced in reviewing books on Goodreads achieve a higher agreement in their evaluations. In line with that reasoning, we find that users who rated more than ten books on Goodreads and additionally wrote more than five freetext reviews (N = 5,464), showed less idiosyncratic variance (18.19% [17.33-19.05]) than the full sample analyzed above. Further, inter-book variance was about twice as large as in the full sample (7.67% [7.31-8.25]), although still smaller than idiosyncratic rater tendencies by a factor of 2.37 ([2.16-2.58]; Nratings = 54,016).
6.2 Analysis of Freetext Reviews
The analysis of practiced reviewers indicates that differences between books can account for at least some rating variance under certain circumstances. However, most of the time, this latent consensus appears to be buried under masses of casual, idiosyncratic evaluations. In that context, it is worthwhile pointing out that Goodreads users do not have to provide proof that they have purchased or read a book before rating it. Thus, our full sample likely includes ratings made for self-presentational purposes, ratings submitted without having read the book, and ratings submitted after a mere skimming of the book contents. Such superficial ratings could have inflated the importance of idiosyncratic reader characteristics and wash out the importance of books in our variance decomposition. Thus, we replicated the analyses above, but focusing on people’s freetext reviews (N=58,199) rather than their star ratings. We assume that written reviews include more reliable information about people’s reading experience as the writing process requires deliberate introspection, and likely decreases the rate of fraudulent reviews (the large majority were written ‘pre-gpt’).
In order to produce the numerical values required for variance decomposition, we condensed each review into a binary sentiment score, using a distilbert model (Hugging Face 2024; Sanh et al. 2019). Given the binary scale of the outcome variable, we updated the previous multilevel model with a logistic link function. Logistic models do not provide the residual variance term required for the computation of ICCs. Thus, we used the latent threshold technique (Devine et al. 2024) under which the binary outcome variable is assumed to have an underlying continuous scale with a logistic error distribution. In our case, the assumption of a latent continuum is warranted given that review sentiments are not required (and in fact unlikely) to be binary by nature and that it is, here, the arbitrary outcome of the sentiment model. Further, the assumption of a logistic error variance does not affect the relative importance of reviewers and books, which came down to a factor of 3.36 [3.08–3.65] in favor of reviewer tendencies. The estimated absolute ICC values ascribed 13.85% [13.22–14.66] of the review variance to differences between reviewers and 4.14% [3.65–4.48] to differences between books.
6.3 Analysis of Review Topics
As a final analysis, we examined whether readers bring up similar issues when reviewing the same book. It might be that readers disagree widely in their overall book ratings, while still perceiving a similar set of strengths and weaknesses per book (of which the importance could vary idiosyncratically). We used a large language model (GPT-4o) to annotate written reviews (N = 26,699) for fiction books regarding the mentioning of feeling bored, addicted, or confused by the book contents, as well as the mentioning of characters, and the author’s unique style or skill in writing. We chose these concepts as they relate to common facets of enjoyment and broad book distinctions (cf., page-turning thrillers vs. literary/character-driven fiction). We also annotated the presence of a book summary in the review as a control variable, as we did not expect books to differ much in whether they elicit summaries from readers. A Python script iterated through the reviews and was given a system prompt to “answer with numerical annotations” pertaining to the themes above (for the full prompt, see finetuning.qmd in supplementary materials). Subsequently, we went through the annotations with an R loop, displaying the LLM responses, and manually scoring the number of mistakes per review between 0 and 6 (based on the number of possible themes; see examine_gpt_performance.R and annotation_finetuning.csv in supplementary materials). GPT-4o achieved an annotation accuracy of 98.6% [97.74-99.11] on 200 human-annotated book reviews.
For our analysis, we drew pairs of reviews either targeting the same book or stemming from the same reviewer. As shown in Figure 3, the presence of specific review attributes was better predicted by other reviews of the same reviewer, compared to other reviews about the same book. Especially the writing of summaries seems to be highly characteristic of specific reviewers.
Figure 3: The X axis shows various topics that were brought up in book reviews. The Y axis shows how much more likely it is that a review mentions this topic if it was mentioned in a reference review, i.e.,
For instance, the black triangular point on the left indicates that a review is twice as likely to mention that the reviewer felt addicted to the book, if the same reviewer made that claim about a different (randomly chosen) book. The gray circle below the triangle indicates that a review is about 1.6 times more likely to contain “feeling addicted” if another (randomly chosen) reviewer made that statement.
7. Discussion
How much value should you place on a stranger’s opinion when looking for an enjoyable read? Our analyses suggest that you should not place any at all. A single book rating will generally tell you much more about the reviewer’s average rating than the book’s – up to ten times more in our analyses.
While it is tempting to proclaim aesthetic value in absolute terms (‘must-reads’ vs. ‘trash-novels’), we find rater characteristics had a much stronger impact on book reviews than a book’s general appeal (here: its average rating). This is in line with previous work on the appreciation of jokes and abstract art (Leder et al. 2016; Rosenbusch et al. 2022). Across star ratings and written reviews, the impact of book characteristics (3-4% of rating variance) was even lower than our a-priori estimates (5-8%).
While genre fans also differed widely in their evaluations of genre books, reviewers with relatively many reviews and ratings achieved higher agreement (7-8%). That is to say that they managed to approximate the average book evaluation better than the average reader, which might be a homogenizing effect of expertise (i.e., a trained eye). However, even for reviews from regular Goodreads review writers, one should consider that their ratings remain virtually uncorrelated to the reading enjoyment of another individual reader. All in all, we conclude that one should not just apply mild but extreme perspectivism when talking about the enjoyability of books.
In addition to the numerical sentiments of written reviews (which likely lost some informative variance during binarization), we also found that specific complaints and compliments are more characteristic of the reader than the book. Thus, we speculate that readers do not just apply a different set of weights for consensual book criteria, but rather apply altogether different criteria (Annalyn et al. 2018; Kraxenberger et al. 2021). This does not mean that book ratings are ‘random’. People read books for different reasons and thus end up bestowing different evaluations and critiques (Kraxenberger et al. 2021). Thus, it is all the more noteworthy when reviewers do agree in their comments, as such an outlying degree of consensus hints at book contents that deviate so much from the norm that readers cannot help but converge in their reviews. For instance, today (30th of September 2024) the latest reviews of the book If on a Winter’s Night a Traveler almost all mention that it is confusing or at least challenging to read (which some readers report to enjoy), indicating that unique books lead to more homogenous review contents. Accordingly, we want to mention that online platforms can increase the informativeness of their comment section by implementing an upvote feature and increase the visibility of highly upvoted comments (as Goodreads does). Such response aggregations can boost the wisdom-of-the-crowd and facilitate faster and more reliable book insights from (shared) reader responses.
In the world of scientific peer review, it is well-known that manuscript evaluations are rarely consensual and often deviate from some ‘true’ metric (Montgomery et al. 2002; Siler et al. 2015). A popular method to strengthen the association between text contents and evaluation is the introduction of narrowly defined rubrics (Jonsson and Svingby 2007). For fiction writing, such rubrics could target criteria which are believed to predict reading enjoyment for a wide audience (e.g., “Does this book feature a protagonist with desirable characteristic X?” or “Is the adverb prevalence in this book below X%?”). While high-dimensional annotations of books can likely be generated by AI models (cf., our annotation of book reviews above, Bail 2024) it remains to be seen whether such rubric-style annotations can approximate a single person’s reading enjoyment, especially as reading preferences clearly vary across, for instance, people’s motivations and personality traits (Annalyn et al. 2018; Kraxenberger et al. 2021).
Viewing the current results, people certainly should not feel discouraged to read reviews or rate books themselves. Such activities are likely to enhance one’s reading appreciation by adding social relevance and offering new perspectives on books (Foasberg 2012).
While absolute book rankings can serve as indicators of prestige (see Feldkamp et al. (2024) for a review of quality metrics), they cannot predict a person’s enjoyment of a book. However, rather than eradicate them, they could simply be seen as descriptive of current audience trends or marketing success. Themed lists and book awards also offer value by directing readers towards specific book contents that they are interested in.
In that regard, it is also noteworthy that the lacking ‘accuracy’ of book evaluations does not necessarily preclude accurate book recommendations. When a reader’s relevant book preferences are measured accurately and matching books can be identified reliably, experts or recommendation engines can still function (as long as the measured preferences are sufficiently stable and highly predictive of reading enjoyment). However, general ratings and rankings can be largely discarded for recommendations and numerical scores of individual readers can be fully discarded, at least for professionally published books.
8. Data Availability
Data can be found here: https://github.com/hannesrosenbusch/are_some_books_better_than_others. It has been archived and is persistently available at: https://doi.org/10.17605/OSF.IO/46H7S.
9. Software Availability
Software can be found here: https://github.com/hannesrosenbusch/are_some_books_better_than_others. It has been archived and is persistently available at: https://doi.org/10.17605/OSF.IO/46H7S.
10. Author Contributions
Hannes Rosenbusch: Conceptualization, Investigation, Methodology, Software, Formal analysis, Validation, Visualization, Writing – original draft
Luke Korthals: Conceptualization, Writing – review & editing
References
Adelson, Jill L., Kathleen M. Cash, Caroline M. Pittard, Christine E. Sherretz, Patrick Pössel, and Allison D. Blackburn (2019). “Measuring Reading Self-Perceptions and Enjoyment: Development and Psychometric Properties of the Reading and Me Survey”. In: Journal of Advanced Academics 30 (3), 355–380. http://doi.org/10.1177/1932202x19843237.
Alharthi, Haifa, Diana Inkpen, and Stan Szpakowicz (2018). “A Survey of Book Recommender Systems”. In: Journal of Intelligent Information Systems 51, 139–160. http://doi.org/10.1007/s10844-017-0489-9.
Annalyn, Ng, Maarten W. Bos, Leonid Sigal, and Boyang Li (2018). “Predicting Personality from Book Preferences with User-Generated Content Labels”. In: IEEE Transactions on Affective Computing 11 (3), 482–492. http://doi.org/10.1109/taffc.2018.2808349.
Bail, Chris. A. (2024). “Can Generative AI Improve Social Science?” In: Proceedings of the National Academy of Sciences of the United States of America 121 (21). http://doi.org/10.1073/pnas.2314021121.
Balée, Susan (2006). “Textual Pleasures and Pet Peeves”. In: The Hudson Review 50 (4), 689–699. https://www.jstor.org/stable/20464508 (visited on 01/06/2026).
Ball, Philip (2004). “Measuring Beauty”. In: Nature. http://doi.org/10.1038/news041011-17.
Barré, Jean, Jean-Baptiste Camps, and Thierry Poibeau (2023). “Operationalizing Canonicity: A quantitative Study of French 19th and 20th Century Literature”. In: Journal of Cultural Analytics 8 (3). http://doi.org/10.22148/001c.88113.
Bates, Douglas, Martin Mächler, Ben Bolker, and Steve Walker (2015). “Fitting Linear Mixed-Effects Models Using lme4”. In: Journal of Statistical Software 67 (1), 1–48. http://doi.org/10.18637/jss.v067.i01.
Bates, Douglas, Martin Maechler, Ben Bolker, Steven Walker, Rune Haubo Bojesen Christensen, Henrik Singmann, Bin Dai, Fabian Scheipl, Gabor Grothendieck, Peter Green, John Fox, Alexander Bauer, Pavel N. Krivitsky, Emi Tanaka, Mikael Jagan, Ross D. Boylan, and Anna Ly (2024). lme4: Linear Mixed-Effects Models using ’Eigen’ and S4. Version 1.1-35.3. [Computer software]. http://doi.org/10.32614/CRAN.package.lme4.
Baumard, Nicolas, Elise Huillery, Alexandre Hyafil, and Liane Safra (2022). “The Cultural Evolution of Love in Literary History”. In: Nature Human Behaviour 6 (4), 506–522. http://doi.org/10.1038/s41562-022-01292-z.
Benkoussas, Chahinez, Hussam Hamdan, Shereen Albitar, Anaïs Ollagnier, and Patrice Bello (2014). “Collaborative Filtering for Book Recommendation”. In: CLEF Working Notes, 501–507. https://ceur-ws.org/Vol-1180/CLEF2014wn-Inex-BenkoussasEt2014.pdf (visited on 01/06/2026).
Bizzoni, Yuri, Ida Marie Lassen, Telma Peura, Mads Rosendahl Thomsen, and Kristoffer Nielbo (2022). “Predicting Literary Quality How Perspectivist Should We Be?” In: Proceedings of the 1st Workshop on Perspectivist Approaches to NLP LREC2022. Ed. by Gavin Abercrombie, Valerio Basile, Sara Tonelli, Verena Rieser, and Alexandra Uma. European Language Resources Association, 20–25. https://aclanthology.org/2022.nlperspectives-1.3/ (visited on 01/06/2026).
Bizzoni, Yuri, Pascale Moreira, Nicole Dwenger, Ida Lassen, Mads Thomsen, and Kristoffer Nielbo (2023). “Good Reads and Easy Novels: Readability and Literary Quality in a Corpus of US-published Fiction”. In: Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). Ed. by Tanel Alumäe and Mark Fishel. University of Tartu Library, 42–51. https://aclanthology.org/2023.nodalida-1.5/ (visited on 01/06/2026).
Boyd, Ryan L., Kate G. Blackburn, and James W. Pennebaker (2020). “The Narrative Arc: Revealing Core Narrative Structures through Text Analysis”. In: Science Advances 6 (32). http://doi.org/10.1126/sciadv.aba2196.
Carlsen, G. Robert and Anne Sherrill (1988). Voices of Readers: How We Come To Love Books. ERIC. https://eric.ed.gov/?id=ED295136 (visited on 01/06/2026).
Chuey, Aaron, Yiwei Luo, and Ellen M. Markman (2024). “Epistemic Language in News Headlines Shapes Readers’ Perceptions of Objectivity”. In: Proceedings of the National Academy of Sciences of the United States of America 121 (20). http://doi.org/10.1073/pnas.2314091121.
Croxson, Pamela L., Lindsey Neeley, and Daniela Schiller (2021). “You Have to Read this”. In: Nature Human Behaviour ( 5), 1466–1468. http://doi.org/10.1038/s41562-021-01221-6.
Devine, Sean, James O. Uanhoro, A. Ross Otto, and Jessica K. Flake (2024). “Approaches for Quantifying the ICC in Multilevel Logistic Models: A Didactic Demonstration”. In: Collabra: Psychology 10 (1). http://doi.org/10.1525/collabra.94263.
Feldkamp, Pascale, Yuri Bizzoni, Mads R. Thomsen, and Kristoffer L. Nielbo (2024). “Measuring Literary Quality. Proxies and Perspectives”. In: Journal of Computational Literary Studies 3 (1). http://doi.org/10.48694/jcls.3908.
Foasberg, Nancy M. (2012). “Online Reading Communities: From Book Clubs to Book Blogs”. In: Journal of Social Media in Society 1 (1). https://thejsms.org/index.php/JSMS/article/view/3 (visited on 01/06/2026).
Ganjigunte Ashok, Vikas, Song Feng, and Yejin Choi (2013). “Success with Style: Using Writing Style to Predict the Success of Novels”. In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Ed. by David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu, and Steven Bethard. Association for Computational Linguistics, 1753–1764. http://doi.org/10.18653/v1/d13-1181.
Gessey-Jones, Thomas, Colm Connaughton, Robin Dunbar, Ralph Kenna, Pádraig MacCarron, Cathal O’Conchobhair, and Joseph Yose (2020). “Narrative Structure of A Song of Ice and Fire Creates a Fictional World with Realistic Measures of Social Complexity”. In: Proceedings of the National Academy of Sciences of the United States of America 117 (46), 28582–28588. http://doi.org/10.1073/pnas.2006465117.
Gordon, Carol and Ya-Ling Lu (2021). “ “I Hate to Read-Or Do I?” Low-Achievers and Their Reading”. In: IASL Conference Proceedings: World Class Learning and Literacy through School Libraries. University of Alberta Libraries. http://doi.org/10.29173/iasl7972.
Greaney, Vincent and Michael Hegarty (1987). “Correlates of Leisure-time Reading”. In: Journal of Research in Reading 10 (1), 3–20. http://doi.org/10.1111/j.1467-9817.1987.tb00278.x.
Hehman, Eric, Clare A. M. Sutherland, Jessica K. Flake, and Michael L. Slepian (2017). “The Unique Contributions of Perceiver and Target Characteristics in Person Perception”. In: Journal of Personality and Social Psychology 113 (4), 513–529. http://doi.org/10.1037/pspa0000090.supp.
Herald, Diana T. and Samuel Stavole-Carter (2019). Genreflecting: A Guide to Popular Reading Interests. Bloomsbury Publishing USA.
Hönekopp, Johannes (2006). “Once more: Is Beauty in the Eye of the Beholder? Relative Contributions of Private and Shared Taste to Judgments of Facial Attractiveness”. In: Journal of Experimental Psychology 32 (2), 199–209. http://doi.org/10.1037/0096-1523.32.2.199.
Hugging Face (2024). Distilbert-base-uncased - finetuned-sst-2-english. https://huggingface.co/distilbert/distilbert-base-uncased-finetuned-sst-2-english. (Visited on 02/10/2026).
Hume, David [1757] (2017). Of the Standard of Taste. Routledge, 483–488.
Hyman, John (2002). “Is Beauty in the Eye of the Beholder?” In: Think 1 (1), 81–92. http://doi.org/10.1017/s1477175600000130.
Jonsson, Anders and Gunilla Svingby (2007). “The Use of Scoring Rubrics: Reliability, Validity and Educational Consequences”. In: Educational Research Review 2 (2), 130–144. http://doi.org/10.1016/j.edurev.2007.05.002.
Kaufman, Scott B. and James C. Kaufman (2007). “Ten Years to Expertise, Many More to Greatness: An Investigation of Modern Writers”. In: The Journal of Creative Behavior 41 (2), 114–124. http://doi.org/10.1002/j.2162-6057.2007.tb01284.x.
King, Stephen (2000). On Writing: A Memoir of the Craft. Simon and Schuster.
Kraaykamp, Gerbert and Katinka Dijkstra (1999). “Preferences in Leisure Time Book Reading: A Study on the Social Differentiation in Book Reading for the Netherlands”. In: Poetics 26 (4), 203–234. http://doi.org/10.1016/s0304-422x(99)00002-9.
Kraxenberger, Maria, Christine A. Knoop, and Winfried Menninghaus (2021). “Who Reads Contemporary Erotic Novels and Why?” In: Humanities and Social Sciences Communications 8, 1–13. http://doi.org/10.1057/s41599-021-00764-3.
Leder, Helmut, Juergen Goller, Tanya Rigotti, and Michael Forster (2016). “Private and Shared Taste in Art and Face Appreciation”. In: Frontiers in Human Neuroscience 10. http://doi.org/10.3389/fnhum.2016.00155.
Lin, Hause, Jana Lasser, Stephan Lewandowsky, Rocky Cole, Andrew Gully, David G Rand, and Gordon Pennycook (2023). “High Level of Correspondence across Different News Domain Quality Rating Sets”. In: PNAS Nexus 2 (9). http://doi.org/10.1093/pnasnexus/pgad286.
Maharjan, Suraj, John Arevalo, Manuel Montes, Fabio A. González, and Thamar Solorio (2017). “A Multi-task Approach to Predict Likability of Books”. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. Ed. by Mirella Lapata, Phil Blunsom, and Alexander Koller. Association for Computational Linguistics, 1217–1227. http://doi.org/10.18653/v1/e17-1114.
Mikkonen, Anna and Pertti Vakkari (2017). “Reader Characteristics, Behavior, and Success in Fiction Book Search”. In: Journal of the Association for Information Science and Technology 68, 2154–2165. http://doi.org/10.1002/asi.23843.
Mittelmark, Howard and Sandra Newman (2009). How Not to Write a Novel: 200 Classic Mistakes and How to Avoid Them–A Misstep-by-misstep Guide. Harper Collins.
Montgomery, Alan A., Anna Graham, Philip H. Evans, and Tom Fahey (2002). “Inter-rater Agreement in the Scoring of Abstracts Submitted to a Primary Care Research Conference”. In: BMC Health Services Research 2 (8), 1–4. http://doi.org/10.1186/1472-6963-2-8.
Moreira, Pascake, Yuri Bizzoni, Kristoffer Nielbo, Ida Lassen, and Mads Thomsen (2023). “Modeling Readers’ Appreciation of Literary Narratives Through Sentiment Arcs and Semantic Profiles”. In: Proceedings of the 5th Workshop on Narrative Understanding. Ed. by Nader Akoury, Elizabeth Clark, Mohit Iyyer, Snigdha Chaturvedi, Faeze Brahman, and Khyathi Chandu. Association for Computational Linguistics, 25–35. http://doi.org/10.18653/v1/2023.wnu-1.5.
Moreira, Pascale and Yuri Bizzoni (2023). “Dimensions of quality: Contrasting stylistic vs. semantic features for modelling literary quality in 9,000 novels”. In: International Conference Recent Advances in Natural Language Processing, RANLP 2023. Ed. by Ruslan Mitkov Galia Angelova Maria Kunilovskaya. Recent Advances in Natural Language Processing. Association for Computational Linguistics. http://doi.org/10.26615/978-954-452-092-2_080.
National Book Foundation (2024). National Book Awards. https://www.nationalbook.org/awards-prizes/national-book-awards-2024/ (visited on 02/10/2026).
Otten, Cord, Michel Clement, and Dominik Stehr (2019). “Sales Estimations in the Book Industry–Comparing Management Predictions with Market Response Models in the Children’s Book Market”. In: Journal of Media Business Studies 16 (4), 249–274. http://doi.org/10.1080/16522354.2019.1623436.
Pelowski, Matthew, Gernot Gerger, Yasmine Chetouani, Patrick S. Markey, and Helmut Leder (2017). “But Is It really Art? The Classification of Images as “Art”/“Not Art”’ and Correlation with Appraisal and Viewer Interpersonal Differences”. In: Frontiers in Psychology 8, 1729. http://doi.org/10.3389/fpsyg.2017.01729.
Poletti, Anna, Judith Seaboyer, Rosanne Kennedy, Tully Barnett, and Kate Douglas (2016). “The Affects of Not Reading: Hating Characters, Being Bored, Feeling Stupid”. In: Arts and Humanities in Higher Education 15 (2), 231–247. http://doi.org/10.1177/1474022214556898.
Ramirez, Gabriel, Laura Fries, Elizabeth Gunderson, Marjorie W. Schaeffer, Erin A. Maloney, Sian L. Beilock, and Susan. C. Levine (2019). “Reading Anxiety: An Early Affective Impediment to Children’s Success in Reading”. In: Journal of Cognitive Development 20 (1), 15–34. http://doi.org/10.1080/15248372.2018.1526175.
Reuter, Kara (2007). “Assessing Aesthetic Relevance: Children’s Book Selection in a Digital Library”. In: Journal of the American Society for Information Science and Technology 58 (12), 1745–1763. http://doi.org/10.1002/asi.20657.
Rosenbusch, Hannes, Anthony M. Evans, and Michael Zeelenberg (2022). “The Relative Importance of Joke and Audience Characteristics in Eliciting Amusement”. In: Psychological Science 33 (9), 1386–1394. http://doi.org/10.1177/09567976221098595.
Sanh, Victor, Lysandre Debut, Julien Chaumond, and Thomas Wolf (2019). “DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter”. In: arXiv preprint. http://doi.org/arXiv.1910.01108.
Schepman, Astrid, Paul Rodway, and Sarah J. Pullen (2015). “Greater Cross-viewer Similarity of Semantic Associations for Representational than for Abstract Artworks”. In: Journal of Vision 15 (14), 12. http://doi.org/10.1167/15.14.12.
Shaffi, Sarah (2023). “James Bond Novels to Be Reissued with Racial References Removed”. In: The Guardian.
Siler, Kyle, Kirby Lee, and Lisa Bero (2015). “Measuring the Effectiveness of Scientific Gatekeeping”. In: Proceedings of the National Academy of Sciences of the United States of America 112 (2), 360–365. http://doi.org/10.1073/pnas.1418218112.
Silva, Giovana D. da, Filipi N. Silva, Henrique F. D. Arruda, Bárbara C. e Souza, Luciano da F. Costa, and Diego R. Amancio (2024). “Using Full-text Content to Characterize and Identify Best Seller Books: A Study of Early 20th-century Literature.” In: PLoS ONE 19 (4). http://doi.org/10.1371/journal.pone.0302070.
Strommen, Linda T. and Barbara F. Mates (2004). “Learning to Love Reading: Interviews With Older Children and Teens”. In: Journal of Adolescent & Adult Literacy 48 (3), 188–200. http://doi.org/10.1598/jaal.48.3.1.
Sujo, Jessie C. Martín, Elisabet Golobardes i Ribé, and Xavier Vilasís Cardona (2021). “CAIT: A Predictive Tool for Supporting the Book Market Operation Using Social Networks”. In: Applied Sciences 12 (1). http://doi.org/10.3390/app12010366.
Tankard, James and Laura Hendrickson (1996). “Specificity, Imagery in Writing: Testing the Effects of “Show, Don’t Tell””. In: Newspaper Research Journal 17 (1–2), 35–48. http://doi.org/10.1177/073953299601700105.
Thelwall, Mike and Kayvan Kousha (2017). “Goodreads: A Social Network Site for Book Readers”. In: Journal of the Association for Information Science and Technology 68 (4), 972–983. http://doi.org/10.1002/asi.23733.
Unrau, Norman J., Robert Rueda, Elena Son, Joshua R. Polanin, Rebecca J. Lundeen, and Alison K. Muraszewski (2018). “Can Reading Self-Efficacy Be Modified? A Meta-Analysis of the Impact of Interventions on Reading Self-Efficacy”. In: Review of Educational Research 88 (2), 167–204. http://doi.org/10.3102/0034654317743199.
Vessel, Edward A., Nataklia Maurer, Alexander H. Denker, and G. Gabrielle Starr (2018). “Stronger Shared Taste for Natural Aesthetic Domains than for Artifacts of Human Culture”. In: Cognition 179, 121–131. http://doi.org/10.1016/j.cognition.2018.06.009.
Walsh, Melanie and Maria Antoniak (2021). “The Goodreads “Classics”: A Computational Study of Readers, Amazon, and Crowdsourced Amateur Criticism”. In: Journal of Cultural Analytics 6 (2). http://doi.org/10.22148/001c.22221.
Wang, Xindi, Burcu Yucesoy, Onur Varol, Tina Eliassi-Rad, and Albert-László Barabási (2019). “Success in Books: Predicting Book Sales before Publication”. In: EPJ Data Science 8 (31), 1–20. http://doi.org/10.1140/epjds/s13688-019-0208-6.
Waples, Douglas (1931). “What Subjects Appeal to the General Reader?” In: Library Quarterly 1 (2), 189–203. http://doi.org/10.1086/612905.
White, Elwyn B. and William Strunk (2023). The Elements of Style. Open Road Media.
Yucesoy, Burcu, Xindi Wang, Junming Huang, and Albert-László Barabási (2018). “Success in Books: A Big Data Approach to Bestsellers”. In: EPJ Data Science 7 (1). http://doi.org/10.1140/EPJDS/S13688-018-0135-Y.


