Pith. sign in

REVIEW 4 major objections 5 minor 97 references

Beyond the Lens: Quantifying the Impact of Scientific Documentaries through Amazon Reviews

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A fine-tuned GPT-4o model can classify how scientific documentaries affect viewers from their Amazon reviews, reaching $F_1$ 0.76 for impact and 0.88 for sentiment, and the method generalizes to new films.

desk verdict A useful new dataset and honest in-sample benchmark, but the paper's generalizability claim rests on a circular evaluation and should be revised before publication. read the letter →

arxiv 2502.08705 v2 pith:6D5TKE73 submitted 2025-02-12 cs.CY cs.DLphysics.ed-ph

classification cs.CYcs.DLphysics.ed-ph
keywords impactanalysisscientificfilmsnaturallanguageprocessingreviewlargemodelsAmazonreviewssentimentdocumentary
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Scientific documentaries are hard to assess at scale: interviews and focus groups give depth but not breadth. This paper argues that ordinary Amazon reviews can fill that gap, and that a fine-tuned large language model can do the reading. The authors build a five-category taxonomy of viewer impact, release 1296 human-annotated sentences from 1043 reviews of six data-driven documentaries, and show that a fine-tuned GPT-4o model reaches $F_1$ scores of 0.76 for impact and 0.88 for sentiment on held-out reviews. They also apply the model to reviews of a seventh documentary and report that its predictions agree with human annotators at about 85 to 88 percent, which they read as evidence the approach generalizes. If right, this gives documentary makers a cheap, reusable way to measure audience engagement and learning signals.

What carries the argument

The load-bearing machinery is the pairing of a five-category impact taxonomy with a sentence-level annotation and classification pipeline. The taxonomy—Shift in Cognition, Attitudes Toward the Film, Interest with Science Topic, Impersonal Report, and Not Applicable—converts open-ended viewer comments into mutually exclusive labels, and the training labels come from majority votes over multiple crowd annotations, with the agreement statistic $\kappa$ ranging from 0.61 to 0.68. The classification design that carries the claim is the prompt that hands a language model the target sentence together with, optionally, the full review as context, followed by fine-tuning of GPT-4o on the training split; the context lets the model interpret sentences whose impact is only clear from surrounding review text.

What would settle it

Have two or three domain experts re-annotate the held-out test sentences with the same taxonomy and compare their labels to both the crowd majority labels and the fine-tuned GPT-4o model's predictions. If expert labels systematically disagree with the crowd labels, especially on 'Impersonal Report' and 'Shift in Cognition', the reported $F_1$ scores would be measuring agreement with a biased gold standard rather than genuine classification ability.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that a large language model can sort viewer sentences into meaningful impact categories, and that adding full-review context and fine-tuning makes the sorting reliable enough to carry over to new films. Using a five-category taxonomy and crowd-annotated sentence labels, the authors compare traditional classifiers, BERT and RoBERTa, several prompt-only LLMs, and a fine-tuned GPT-4o model. The fine-tuned model performs best, with $F_1$ 0.76 for impact and 0.88 for sentiment on the held-out test set; 'Attitudes Toward the Film' is predicted best ($F_1$ 0.86) and 'Impersonal Report' worst ($F_1$ 0.42), mirroring the lower human agreement on that class. On 400 sentences from reviews of the Hubble documentary, the model's impact labels matched two human annotators in about 85.3 percent of cases and sentiment in about 87.6 percent, which the paper takes as evidence of generalizability beyond the six training films.

Load-bearing premise

The load-bearing premise is that the majority-vote labels from volunteer annotators, whose agreement was moderate rather than strong, are correct enough to serve as the gold standard for training and evaluating the classifiers.

Editorial extensions

If this is right

  • Documentary production teams can use the released dataset and fine-tuned classifier to get large-scale audience-impact feedback from Amazon reviews, complementing small qualitative studies.
  • The classifier can be applied to other cinematic scientific-visualization documentaries beyond the six training films, with the Hubble review holdout serving as the paper's evidence.
  • Adding the full review as context improves sentiment classification and helps impact classification to a lesser degree, while fine-tuning produces the largest performance gain.
  • The five-category taxonomy offers a reusable scheme for coding viewer engagement and learning signals in review text.
  • The thematic patterns—for example, positive attitudes cite engagement and informativeness while negative attitudes cite narration or production quality—give content creators concrete levers for what drives viewer reactions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the model's worst category, Impersonal Report, is also the category with the lowest human agreement, part of the reported $F_1$ gap may be label noise rather than model failure; a useful extension would be a two-stage classifier that first separates descriptive from evaluative sentences.
  • The paper stops short of testing cross-platform transfer; running the same fine-tuned model on YouTube comments for the same films, which the authors note skew negative, would directly test whether the classifier's accuracy is platform-dependent.
  • The taxonomy collapses 'shift in knowledge' and 'disagreement with the science' into one category, so a viewer who reports learning and a viewer who rejects the film's science are treated as the same impact; future work might split these to distinguish education from resistance.
  • Because the reviewed sentences are short, averaging about 11 words, the method may not transfer to long-form critical reviews without adaptation even within English.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces a sentence-level annotated dataset of 1,286 Amazon reviews across six scientific documentaries, a five-category impact taxonomy adapted from prior work, and an evaluation of SVM, logistic regression, decision trees, BERT, RoBERTa, and several LLMs (GPT-3.5, GPT-4, GPT-4o, Llama 3, Mixtral) for impact and sentiment classification. The best model, a fine-tuned GPT-4o, reaches test F1 of 0.76 for impact and 0.88 for sentiment on a held-out 30% split. The paper also reports a thematic analysis of impact categories and a 'generalizability test' on 400 sentences from reviews of the Hubble documentary.

Significance. The annotated dataset and the detailed comparison of classifiers are useful resources for studying public engagement with scientific documentaries. The in-sample evaluation is largely sound: a held-out test set is used, multiple model families are compared, and the prompt design is documented. The main claim that the best classifier is 'generalizable to other datasets' is not established, because the Section 5.3 test lacks independent gold labels. If that gap is repaired, the contribution would be a practical tool for documentary producers and a reusable benchmark for future work.

major comments (4)
  1. [5.3] The generalizability test does not measure out-of-domain classification accuracy. The model is applied to 400 Hubble sentences, but no independent gold labels are produced for this corpus; instead, two annotators are asked to evaluate the model's predicted labels, yielding mean agreement of 85.3% for Impact and 87.6% for Sentiment. Because the annotators see the predictions they are judging, this protocol can confirm plausible labels without measuring correctness, and no inter-annotator reliability, annotation protocol, or blinding is reported. The conclusion in Section 7 that the model 'should be applicable to a wide range of scientific CSV-style documentaries' therefore rests on invalid evidence. I recommend annotating the Hubble sentences with the same majority-vote protocol used for the training data, or using another independently labeled held-out corpus.
  2. [Table 5] The per-class F1 scores weaken the claim in Section 5 that the models 'successfully identify' Impact categories. Impersonal Report has F1 of 0.42, and the confusion matrix in Table 4 contains only 12 true instances for this class (the annotation distribution in Figure 1 also shows it is rare). The paper acknowledges the imbalance but does not address the consequence that the classifier is effectively unreliable for this category, which is one of the five categories in the taxonomy. A per-class confidence interval or a threshold-based discussion of acceptable performance for rare classes is needed before the overall F1 of 0.76 can be interpreted as successful.
  3. [Table 3] The reported F1 differences between models, such as 0.76 for fine-tuned GPT-4o versus 0.71 for GPT-4o with context, are presented without confidence intervals or significance tests. With a test set of roughly 386 sentences, differences of this magnitude may be within sampling noise, so the claim that fine-tuning 'confirms usefulness' (Section 5.1) is not statistically supported. I recommend bootstrap confidence intervals or paired significance tests across the test set.
  4. [Appendix B, Figure 3] The prompt used for LLM classification has different label names from the taxonomy in Table 1: 'Engagement with Film' instead of 'Attitudes Toward the Film' and 'Shift in Knowledge' instead of 'Shift in Cognition'. This mismatch makes the experimental setup hard to reproduce and could affect LLM predictions if the models rely on label semantics. Please align the prompt labels with the taxonomy or explain the mapping.
minor comments (5)
  1. [3.2] The taxonomy is described as 'novel' but is adapted from Rezapour and Diesner [75]; please describe the specific modifications more clearly.
  2. [Table 3] The table formatting is inconsistent in places, such as the GPT-4 w/o context sentiment recall appearing as '73' rather than '0.73'.
  3. [Table 4] The row sums of the confusion matrix do not match the expected test set size (rows sum to 366 instead of 386); please verify the counts.
  4. [5.3] Please specify who the two annotators are, whether they are authors or external, and whether they were blind to the purpose of the evaluation.
  5. [Figure 1] The caption says 'full 1286 sentences'; consider changing to 'final 1286 sentences' for clarity.

Circularity Check

1 steps flagged · score 6.0 of 10

Generalizability evidence is self-referential: annotators rate the model's own predicted labels, so the out-of-domain claim does not rest on an independent gold standard.

  1. other [Section 5.3 (Generalizability Test), with the conclusion restated in Section 7]
    "To assess the model's performance, we tasked two annotators with evaluating the accuracy of the predicted labels for both the Impact and Sentiment categories. The mean agreement for Impact is approximately 85.3%. For Sentiment agreement is∼87.6%, demonstrating a substantial level of alignment between the annotators and the classification model."

    The generalizability test provides no independent gold-standard labels for the 400 Hubble sentences. The model first predicts Impact and Sentiment labels, and then two annotators are asked to evaluate the accuracy of those predicted labels; the reported 85.3%/87.6% agreement is a measure of how plausible the model's own outputs look to raters, not a comparison against an external ground truth. Because the evaluation criterion is the model output being evaluated, the Section 7 conclusion that 'the model should be applicable to a wide range of scientific CSV-style documentaries' reduces to a self-referential check.

full rationale

The main supervised classification pipeline is not circular: the fine-tuned GPT4o model is trained on a 70% split of the Zooniverse-labeled sentences and evaluated on the held-out 30% test set (Sections 3.3, 3.4, 4.1, Table 3), so the reported 0.76 Impact / 0.88 Sentiment F1 is an independent, in-sample measurement against human majority-vote labels. The taxonomy is transparently adapted from the authors' own prior work [75]; that self-citation is disclosed and does not by itself force any result. The circular content is concentrated in the generalization claim. Section 5.3 does not produce Hubble gold labels; it has two annotators rate the model's own predictions and reports the resulting agreement as evidence of generalization, and Section 7 then treats this as showing applicability to a wide range of CSV documentaries. That is a self-referential evaluation, not an external benchmark. The score of 6 reflects partial circularity: the in-sample result stands, but the out-of-domain half of the central claim is supported only by a check of the model against its own outputs.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests primarily on the quality of the annotated dataset and the evaluation design; no new physical or mathematical entities are introduced.

free parameters (2)
  • Fine-tuning hyperparameters = epochs=3, batch_size=8, warmup_steps=500, weight_decay=0.01
    Chosen without a systematic search; affects GPT4o fine-tuned performance (F1 0.76 impact, 0.88 sentiment).
  • LLM decoding temperature = 0
    Set for consistency; not varied, so the effect on scores is unknown.
assumptions (4)
  • domain assumption Amazon review sentences are a valid proxy for viewer impact of documentaries.
    Data collection Section 3.1; relies on self-reported, single-time comments, as acknowledged in Section 8.
  • domain assumption Impact categories are mutually exclusive and sentence-level majority vote provides a reliable gold standard.
    Annotation procedure Sections 3.3-3.4; Cohen's Kappa 0.61-0.68 shows moderate/substantial agreement, with lower reliability for rare classes.
  • domain assumption The model's test-set performance transfers to other documentaries.
    Generalizability test Section 5.3 uses only one additional film (Hubble) and annotator assessment of model labels rather than independent gold labels.
  • standard math Standard fine-tuning of pretrained transformers is an appropriate inductive bias for this classification task.
    Methodology Section 4.1; standard practice, not demonstrated on this domain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond the Lens: Quantifying the Impact of Scientific Documentaries through Amazon Reviews." pith.science (2026). https://pith.science/paper/6D5TKE73

@misc{pith2026250208705,
  author       = {Pith},
  title        = {Pith review of: Beyond the Lens: Quantifying the Impact of Scientific Documentaries through Amazon Reviews},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6D5TKE73}},
  note         = {Machine review of arXiv:2502.08705}
}
read the original abstract

Engaging the public with science is critical for a well-informed population. A popular method of scientific communication is documentaries. Once released, it can be difficult to assess the impact of such works on a large scale, due to the overhead required for in-depth audience feedback studies. In what follows, we overview our complementary approach to qualitative studies through quantitative impact and sentiment analysis of Amazon reviews for several scientific documentaries. In addition to developing a novel impact category taxonomy for this analysis, we release a dataset containing 1296 human-annotated sentences from 1043 Amazon reviews for six movies created in whole or part by the Advanced Visualization Lab (AVL). This interdisciplinary team is housed at the National Center for Supercomputing Applications and consists of visualization designers who focus on cinematic presentations of scientific data. Using this data, we train and evaluate several machine learning and large language models, discussing their effectiveness and possible generalizability for documentaries beyond those focused on for this work. Themes are also extracted from our annotated dataset which, along with our large language model analysis, demonstrate a measure of the ability of scientific documentaries to engage with the public.

Figures

Figures reproduced from arXiv: 2502.08705 by the authors.

Figure 1
Figure 1. Breakdown of full 1286 sentences in the annotated [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. First stage of annotation process in which users are instructed to select their first choice for the sentiment category [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Prompt used for prediction of impact categories [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

97 extracted references · 71 canonical work pages

  1. [1]

    Valletta, and Alex Mesoudi

    Alberto Acerbi, Josh Burns, Umut Cabuk, Jakub Kryczka, Bettina Trapp, John J. Valletta, and Alex Mesoudi. 2023. Sentiment analysis of the Twitter response to Netflix’s Our Planet documentary.Conservation Biology: The Journal of the Society for Conservation Biology 37, 4 (2023), e14060. https://doi.org/10.1111/cobi.14060

  2. [2]

    AI@Meta. 2024. Llama 3 Model Card. (2024). https://github.com/meta-llama/ llama3/blob/main/MODEL_CARD.md

  3. [3]

    Dimosthenis Antypas, Alun Preece, and Jose Camacho-Collados. 2023. Negativity spreads faster: A large-scale multilingual twitter analysis on the role of sentiment in political communication. Online Social Networks and Media 33 (2023), 100242. https://doi.org/10.1016/j.osnem.2023.100242

  4. [4]

    Orestes Appel, Francisco Chiclana, Jenny Carter, and Hamido Fujita. 2016. A hy- brid approach to the sentiment analysis problem at the sentence level.Knowledge- Based Systems 108 (2016), 110–124. https://doi.org/10.1016/j.knosys.2016.05.040 New Avenues in Knowledge Bases for Natural Language Processing

  5. [5]

    Fatemeh Ardestani. 2024. YouTube emotions and ratings on Amazon: An in-depth analysis. (2024)

  6. [6]

    Fernando Arias, Mayteé Zambrano Núñez, Ariel Guerra-Adames, Nathalia Tejedor-Flores, and Miguel Vargas-Lombardo. 2022. Sentiment Analysis of Pub- lic Social Media as a Tool for Health-Related Topics. IEEE Access 10 (2022), 74850–74872. https://doi.org/10.1109/ACCESS.2022.3187406

  7. [7]

    al Arif Ridho Lubis

    et. al Arif Ridho Lubis. 2023. Comparison of Transformer Based and Traditional Models on Sentiment Analysis on Social Media Datasets. IEEE Xplore (2023). https://ieeexplore.ieee.org/document/10331232

  8. [8]

    Agnaldo Arroio. 2010. Context based learning: A role for cinema in science education. Science Education International 21 (09 2010)

Show all 97 references
  1. [9]

    Eylem Atakav. 2024. The impact of documentary filmmaking. Academic Quarter| Akademisk kvarter (2024)

  2. [10]

    Steve Olusegun Bada and Steve Olusegun. 2015. Constructivism learning theory: A paradigm for teaching and learning. Journal of Research & Method in Education 5, 6 (2015), 66–70

  3. [11]

    Malti Bansal, Apoorva Goyal, and Apoorva Choudhary. 2022. A comparative analysis of K-Nearest Neighbor, Genetic, Support Vector Machine, Decision Tree, and Long Short Term Memory algorithms in machine learning.Decision Analytics Journal 3 (2022), 100071. https://doi.org/10.101...

  4. [12]

    Barnes and Christopher J

    David G. Barnes and Christopher J. Fluke. 2008. Incorporating interactive three- dimensional graphics in astronomy research papers.NA 13, 8 (Nov. 2008), 599–605. https://doi.org/10.1016/j.newast.2008.03.008 arXiv:0709.2734 [astro-ph]

  5. [13]

    Diana Barrett and Sheila Leddy. 2008. Assessing creative media’s social impact . Fledgling Fund Wilmington, DE, USA

  6. [14]

    Ashley Bieniek-Tobasco, Sabrina McCormick, Rajiv N Rimal, Cherise B Harring- ton, Madelyn Shafer, and Hina Shaikh. 2019. Communicating climate change through documentary film: Imagery, emotion, and efficacy. Climatic Change 154 (2019), 1–18

  7. [15]

    Kalina Borkiewicz, Eric Jensen, Stuart Levy, and Jill P Naiman. 2022. Introduc- ing cinematic scientific visualization: a new frontier in science communication. Impact of Social Sciences Blog (2022)

  8. [16]

    Kalina Borkiewicz, Jill P Naiman, and Haoming Lai. 2019. Cinematic visualization of multiresolution data: Ytini for adaptive mesh refinement in houdini. The Astronomical Journal 158, 1 (2019), 10

  9. [17]

    Kalina Borkiewicz, J. P. Naiman, and Haoming Lai. 2019. Cinematic Visualization of Multiresolution Data: Ytini for Adaptive Mesh Refinement in Houdini. Astron. J. 158, 1, Article 10 (July 2019), 10 pages. https://doi.org/10.3847/1538-3881/ab1f6f arXiv:1808.02860 [cs.GR]

  10. [18]

    Layla Bouzoubaa and Rezvaneh Rezapour. 2024. Euphoria’s Hidden Voices: Ex- amining Emotional Resonance and Shared Substance Use Experience of Viewers on Reddit

  11. [19]

    Zoë Buck. 2013. The effect of color choice on learner interpretation of a cosmology visualization. Astronomy Education Review 12, 1 (2013)

  12. [20]

    Nick Cawthon and Andrew Vande Moere. 2007. The effect of aesthetic on the usability of data visualization. In 2007 11th International Conference Information Visualization (IV’07). IEEE, 637–648

  13. [21]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al . 2024. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology 15, 3 (2024), 1–45

  14. [22]

    Chaomei Chen. 2005. Top 10 unsolved information visualization problems. IEEE computer graphics and applications 25, 4 (2005), 12–16

  15. [23]

    Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1 (1960), 37–46

  16. [24]

    David Conrad-Pérez, Caty Borum, Jacqueline Olive, Lisa Flick Wilson, Vanessa Jackson, and Shakita Brooks Jones. 2022. Breaking cultures of silence: Learnings from a participatory community-centred approach to leveraging and researching documentaries for social change. Journal ...

  17. [25]

    John Corner. 2002. Performing the real: Documentary diversions. Television & new media 3, 3 (2002), 255–269

  18. [26]

    Katsiaryna Cortis and Brian Davis. 2021. Over a decade of social opinion mining: a systematic review. Artificial Intelligence Review 54 (2021), 4873–4965. https: //doi.org/10.1007/s10462-021-10030-2

  19. [27]

    Michael F Dahlstrom. 2014. Using narratives and storytelling to communicate science with nonexpert audiences. Proceedings of the national academy of sciences 111, supplement_4 (2014), 13614–13620

  20. [28]

    Xiang Deng, Vasilisa Bashlovkina, Feng Han, Simon Baumgartner, and Michael Bendersky. 2023. What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis. In Companion Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (WWW ’23 Companion...

  21. [29]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR abs/1810.04805 (2018). arXiv:1810.04805 http://arxiv.org/abs/1810.04805

  22. [30]

    Jana Diesner and Rezvaneh Rezapour. 2015. Social computing for impact as- sessment of social change projects. In Social Computing, Behavioral-Cultural Modeling, and Prediction: 8th International Conference, SBP 2015, Washington, DC, USA, March 31-April 3, 2015. Proceedings 8 ....

  23. [31]

    Jana Diesner, Rezvaneh Rezapour, and Ming Jiang. 2016. Assessing public aware- ness of social justice documentary films based on news coverage versus social media. IConference 2016 Proceedings (2016)

  24. [32]

    Leroy W Dubeck, Suzanne E Moshier, and Judith E Boss. 2004. Fantastic voyages: Learning science through science fiction films. Springer Science & Business Media

  25. [33]

    Imane El Alaoui, Younès Gahi, Rachid Messoussi, Youssef Chaabi, Annick To- doskoff, and Abdelilah Kobi. 2018. A novel adaptable approach for senti- ment analysis on big social data. Journal of Big Data 5, 1 (2018), 12. https: //doi.org/10.1186/s40537-018-0120-0

  26. [34]

    Steven L Franconeri, Lace M Padilla, Priti Shah, Jeffrey M Zacks, and Jessica Hull- man. 2021. The science of visual data communication: What works. Psychological Science in the public interest 22, 3 (2021), 110–161

  27. [35]

    John Fraser, Joe E Heimlich, John Jacobsen, Victor Yocco, Jessica Sickler, Jim Kisiel, Mary Nucci, Lance Ford Jones, and Jeanie Stahl. 2012. Giant screen film and science learning in museums. Museum Management and Curatorship 27, 2 (2012), 179–195

  28. [36]

    Sunanda Prabhu Gaunkar, Ellen Askey, Meira Chasman, Kosuke Takaira, Calahan Smith, Amanda Murphy, and Nancy Kawalek. 2022. Exploring the effectiveness of documentary film for science communication. In 2022 IEEE International Con- ference on Quantum Computing and Engineering (Q...

  29. [37]

    Patrick Gerard, Nicholas Botzer, and Tim Weninger. 2023. Truth Social Dataset. Proceedings of the International AAAI Conference on Web and Social Media 1 (Jun. 2023), 1034–1040

  30. [38]

    Daniel Ginting, Ross M Woods, Yusawinur Barella, Liem Satya Limanta, Ahmad Madkur, and Heng Ee How. 2024. The Effects of Digital Storytelling on the Retention and Transferability of Student Knowledge. SAGE Open 14, 3 (2024), 21582440241271267

  31. [39]

    Jane Gregory and Steve Miller. 1998. Science in public: Communication, culture, and credibility. Plenum Press

  32. [40]

    Alessio Guerra and Oktay Karakuş. 2023. Sentiment analysis for measuring hope and fear from Reddit posts during the 2022 Russo-Ukrainian conflict. Frontiers in Artificial Intelligence 6 (2023). https://doi.org/10.3389/frai.2023.1163577 Websci ’25, May 20–24, 2025, New Brunswic...

  33. [41]

    Karen Hirsch and Matt Nisbet. 2007. Documentaries on a Mission: How nonprofits are making movies for public engagement. Future of Public Media Project Re- port. A vailable online: https://cmsimpact. org/resource/documentaries-on-a-mission- how-nonprofitsare-making-movies-for-p...

  34. [42]

    Shaik Asif Hussain and Sana Al Ghawi. 2023. Sentiment Analysis of Real-Time Health Care Twitter Data Using Hadoop Ecosystem. In Hybrid Intelligent Sys- tems, Ajith Abraham, Tzung-Pei Hong, Ketan Kotecha, Kun Ma, Pooja Manghir- malani Mishra, and Niketa Gandhi (Eds.). Springer ...

  35. [43]

    Eric Allen Jensen. 2017. Putting the methodological brakes on claims to measure national happiness through Twitter: Methodological limitations in social media analytics. PLOS ONE 12, 9 (Sept. 2017), e0180080. https://doi.org/10.1371/journal. pone.0180080 Publisher: Public Libr...

  36. [44]

    Jensen, Kalina Borkiewicz, Jeff Carpenter, Stuart Levy, and Jill P

    Eric A. Jensen, Kalina Borkiewicz, Jeff Carpenter, Stuart Levy, and Jill P. Naiman

  37. [45]

    Eric A Jensen, Kalina Borkiewicz, Jill P Naiman, Stuart Levy, and Jeff Carpenter

  38. [46]

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al . 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024)

  39. [47]

    Keshav Kapur and Rajitha Harikrishnan. 2022. Comparative Study of Sentiment Analysis for Multi-Sourced Social Media Platforms. arXiv:2212.04688 [cs.CL] https://arxiv.org/abs/2212.04688

  40. [48]

    al Kian Long Tan, Chin Poo Lee

    et. al Kian Long Tan, Chin Poo Lee. 2022. RoBERTa-LSTM: A Hybrid Model for Sentiment Analysis With Transformer and Recurrent Neural Network. IEEE Xplore (2022). https://ieeexplore.ieee.org/document/9716923

  41. [49]

    Michelle S Lam, Janice Teoh, James A Landay, Jeffrey Heer, and Michael S Bern- stein. 2024. Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooM. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–28

  42. [50]

    Elsie Lee-Robbins and Eytan Adar. 2022. Affective learning objectives for commu- nicative visualizations. IEEE Transactions on Visualization and Computer Graphics 29, 1 (2022), 1–11

  43. [51]

    Elsie Lee-Robbins and Eytan Adar. 2022. Affective Learning Objectives for Communicative Visualizations. http://arxiv.org/abs/2208.04078 arXiv:2208.04078 [cs]

  44. [52]

    Bing Liu and Lei Zhang. 2012. A Survey of Opinion Mining and Sentiment Analysis. Springer US, 415–463. https://doi.org/10.1007/978-1-4614-3223-4_13

  45. [53]

    Bing Liu and Lei Zhang. 2012. A Survey of Opinion Mining and Sentiment Analysis. Springer US, Boston, MA, 415–463. https://doi.org/10.1007/978-1-4614-3223- 4_13

  46. [54]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/1907.11692

  47. [55]

    Like, comment, and share

    Tianhua Luo, Christopher Freeman, and Jillian Stefaniak. 2020. "Like, comment, and share"—professional development through social media in higher education: A systematic review. Education Tech Research Dev 68 (2020), 1659–1683. https: //doi.org/10.1007/s11423-020-09790-5

  48. [56]

    De La Cruz Lyberius Ennio F

    Arvin R. De La Cruz Lyberius Ennio F. Taruc. 2024. Quantifying the Effectiveness of Student Organization Activities using Natural Language Processing. arXiv preprint arXiv:2408.08694 (2024). https://arxiv.org/abs/2408.08694

  49. [57]

    MagellanTV

    Advanced Visualization Lab (NCSA). MagellanTV. 2015. Solar Superstorms: Journey to the Center of the Sun

  50. [58]

    MagellanTV

    Advanced Visualization Lab (NCSA). MagellanTV. 2015. SuperTornado: Anatomy of a MegaDisaster

  51. [59]

    MagellanTV

    Advanced Visualization Lab (NCSA). MagellanTV. 2017. Seeing the Beginning of Time

  52. [60]

    MagellanTV

    Advanced Visualization Lab (NCSA). MagellanTV. 2018. Space Junk

  53. [61]

    MagellanTV

    Advanced Visualization Lab (NCSA). MagellanTV. 2019. Birth of Planet Earth

  54. [62]

    R Malini and MR Sunitha. 2019. Opinion mining on movie reviews. In 2019 1st International Conference on Advances in Information Technology (ICAIT) . IEEE, 282–286

  55. [63]

    Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22, 3 (2012), 276–282

  56. [64]

    Soufiane El Mrabti, Jaouad El-Mekkaoui, Adil Hachmoud, and Mohamed Lazaar

  57. [65]

    Junichiro Niimi. 2024. Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evalua- tion. arXiv preprint arXiv:2407.13069 (2024). https://arxiv.org/abs/2407.13069

  58. [66]

    Matthew C Nisbet and Patricia Aufderheide. 2009. Documentary film: Towards a research agenda on forms, functions, and impacts. Mass Communication and Society 12, 4 (2009), 450–456

  59. [67]

    Knowledge-Based Systems (2024)

    An explainable machine learning model for sentiment analysis of online reviews. Knowledge-Based Systems (2024). https://api.semanticscholar.org/ CorpusID:271804152

  60. [68]

    OpenAI. 2023. ChatGPT: Optimizing Language Models for Dialogue. https: //openai.com/chatgpt. Accessed: 2024-11-26

  61. [69]

    OpenAI. 2023. GPT-4 Technical Report. https://openai.com/research/gpt-4. Accessed: 2024-11-26

  62. [70]

    National Academies of Sciences Engineering and Medicine. 2017. Communi- cating Science Effectively: A Research Agenda . The National Academies Press, Washington, DC. https://doi.org/10.17226/23674

  63. [71]

    Punzo, J

    D. Punzo, J. M. van der Hulst, J. B. T. M. Roerdink, T. A. Oosterloo, M. Ramatsoku, and M. A. W. Verheijen. 2015. The role of 3-D interactive visualization in blind surveys of H I in galaxies. Astronomy and Computing 12 (Sept. 2015), 86–99. https://doi.org/10.1016/j.ascom.2015...

  64. [72]

    Cheng Qian, Nitya Mathur, Nor Hidayati Zakaria, Rameshwar Arora, Vedika Gupta, and Mazlan Ali. 2022. Understanding public opinions on social media for financial sentiment analysis using AI-based techniques. Information Processing & Management 59, 6 (2022), 103098. https://doi....

  65. [73]

    Neelam Chaplot Palak Baid, Apoorva Gupta. 2017. Sentiment Analysis of Movie Reviews using Machine Learning Techniques. International Journal of Computer Applications 179, 7 (Dec 2017), 45–49. https://doi.org/10.5120/ijca2017916005

  66. [74]

    Ana Reyes-Menendez, José Ramón Saura, and Cesar Alvarez-Alonso. 2018. Un- derstanding #WorldEnvironmentDay User Opinions in Twitter: A Topic-Based Sentiment Analysis Approach. International Journal of Environmental Research and Public Health 15, 11 (2018). https://doi.org/10.3...

  67. [75]

    Rezvaneh Rezapour and Jana Diesner. 2017. Classification and detection of micro- level impact of issue-focused documentary films based on reviews. InProceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. 1419–1431

  68. [76]

    al Ravita Chahar

    et. al Ravita Chahar. 2023. Machine Learning-based Sentiment Analysis: A Com- prehensive Review. IEEE Xplore (2023). https://ieeexplore.ieee.org/document/ 10533989

  69. [77]

    RLJ Entertainment

    Advanced Visualization Lab (NCSA). RLJ Entertainment. 2012. Space Junk

  70. [78]

    Ellen Agustine Saputra. 2022. Impactful Storytelling and Social Advocacy in Doc- umentary Filmmaking: Studies of Documentary Impact Methods. In International Conference on Sustainability in Creative Industries . Springer, 165–173

  71. [79]

    Rezvaneh Rezapour, Lufan Wang, Omid Abdar, and Jana Diesner. 2017. Iden- tifying the overlap between election result and candidates’ ranking based on hashtag-enhanced, lexicon-based sentiment analysis. In 2017 IEEE 11th Interna- tional Conference on Semantic Computing (ICSC) ....

  72. [80]

    Lisa Smith, Kimberly Arcand, Randall Smith, Jay Bookbinder, and Jeffrey Smith

  73. [81]

    Lisa F Smith, Jeffrey K Smith, Kimberly K Arcand, Randall K Smith, and Jay A Bookbinder. 2015. Aesthetics and Astronomy: How Museum Labels Affect the Understanding and Appreciation of Deep-Space Images. Curator: The Museum Journal 58, 3 (2015), 282–297

  74. [82]

    Md Shohel Sayeed, Varsha Mohan, and Kalaiarasi Sonai Muthu. 2023. Bert: A review of applications in sentiment analysis. HighTech and innovation journal 4, 2 (2023), 453–462

  75. [83]

    TVF International

    Advanced Visualization Lab (NCSA). TVF International. 2013. Solar Superstorms

  76. [84]

    Frédéric P. A. Vogt, Chris I. Owen, Lourdes Verdes-Montenegro, and Sanchayeeta Borthakur. 2016. Advanced Data Visualization in Astrophysics: The X3D Pathway. ApJ 818, 2, Article 115 (Feb. 2016), 115 pages. https://doi.org/10.3847/0004- 637X/818/2/115 arXiv:1510.02796 [astro-ph.IM]

  77. [85]

    Samer Abdulateef Waheeb, Naseer Ahmed Khan, and Xuequn Shang. 2022. Topic Modeling and Sentiment Analysis of Online Education in the COVID-19 Era Using Social Networks Based Datasets. Electronics 11, 5 (2022). https://doi.org/ 10.3390/electronics11050715

  78. [86]

    Toni Myers

    Advanced Visualization Lab (NCSA). Toni Myers. 2010. Hubble

  79. [87]

    Keith Woodward, John Paul Jones III, Linda Vigdor, Sallie A Marston, Harriet Hawkins, and Deborah P Dixon. 2015. One sinister hurricane: Simondon and collaborative visualization. Annals of the Association of American Geographers 105, 3 (2015), 496–511

  80. [88]

    Yang Yang and Jill E Hobbs. 2020. The power of stories: Narratives and informa- tion framing effects in science communication. American Journal of Agricultural Economics 102, 4 (2020), 1271–1296

  81. [89]

    Moran Yarchi, Christian Baden, and Neta Kligler-Vilenchik. 2021. Political Polar- ization on the Digital Sphere: A Cross-platform, Over-time Analysis of Interac- tional, Positional, and Affective Polarization on Social Media. Political Communi- cation 38, 1–2 (2021), 98–139. h...

  82. [90]

    David Whiteman. 2009. Documentary film as policy analysis: The impact of yes, in my backyard on activists, agendas, and policy. Mass Communication and Society 12, 4 (2009), 457–477

  83. [91]

    Sara K Yeo, Andrew R Binder, Michael F Dahlstrom, and Dominique Brossard

  84. [92]

    first choice

    Liang Yue, Wen Chen, Xiaoyu Li, Weiliang Zuo, and Minghao Yin. 2019. A survey of sentiment analysis in social media. Knowledge and Information Systems 60 (2019), 617–663. https://doi.org/10.1007/s10115-018-1236-4 A Annotation Platform We used the Zooniverse for the annotation ...

  85. [94]

    al Yasir Ali Solangi

    et. al Yasir Ali Solangi. 2019. Review on Natural Language Processing (NLP) and Its Toolkits for Opinion Mining and Sentiment Analysis. IEEE Xplore (2019). https://ieeexplore.ieee.org/abstract/document/8629198

  86. [2017]

    Journal of Science Communication 16, 5 (2017), A02

    Capturing the many faces of an exploded star: communicating complex and evolving astronomical data. Journal of Science Communication 16, 5 (2017), A02

  87. [2018]

    Journal of Science Commu- nication 17, 2 (2018), A07

    An inconvenient source? Attributes of science documentaries and their Beyond the Lens Websci ’25, May 20–24, 2025, New Brunswick, NJ, USA effects on information-related behavioral intentions. Journal of Science Commu- nication 17, 2 (2018), A07

  88. [2023]

    Sustainability 15, 8 (2023), 6845

    Evidence-Based Methods of Communicating Science to the Public through Data Visualization. Sustainability 15, 8 (2023), 6845

  89. [2024]

    Transforming Science Narratives? The Impact of Explanatory Labels of 3D Data Visualization on Public Understanding of Space Science. (2024)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.