Pith. sign in

REVIEW 3 major objections 6 minor 37 references

Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Adding LLM-generated social context to excerpts from long conversations makes them significantly easier to understand and more empathetic.

desk verdict Useful dataset and a plausible result, but the evaluation has a length confound and an internal contradiction in the reported numbers that needs fixing before the mechanism claim can be trusted. read the letter →

arxiv 2412.19966 v1 pith:CZ5WNLBT submitted 2024-12-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords effectivecontextualizationsalientexcerptslong-formsocialconversationslargelanguagemodelscomprehensionempathyconversationsummarizationHSEdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a short, LLM-generated context block placed around a human-highlighted excerpt from a long social conversation significantly improves reader comprehension, readability, and empathy for the speaker. It introduces two contextualization strategies: implicit zero-shot contextualization and explicit prompting for socially grounded attributes such as speaker characteristics, motivation, and relevant locations or events. Both contextualized versions receive significantly higher human ratings than the raw excerpts alone, with the explicit strategy generally rated better on readability, low redundancy, and empathy. The paper also shows that summaries built from context-enriched excerpts are rated more informative and coherent than summaries of the full conversation or of excerpts alone. A new Human-annotated Salient Excerpts (HSE) dataset is released to support further work.

What carries the argument

The central device is the Context-Enriched Excerpt (CEE), produced by prompting an LLM with the full conversation transcript plus a human-highlighted excerpt. The explicit variant, CEEe, uses in-context learning to instruct the model to surface speaker characteristics, speaker motivation, relevant locations or events, and supporting verbatim sentences, and to justify each addition; the implicit variant, CEEi, relies on zero-shot instruction to contextualize without naming social attributes. This contrast is what lets the paper attribute improved comprehension to socially grounded contextualization rather than to generic elaboration.

What would settle it

Show evaluators the same excerpt paired with two equally long context blocks, one containing the speaker's social attributes and one containing matched-length generic background, and measure comprehension with free-response questions; if the generic block scores as high, the claim that social contextualization drives the improvement is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that effective contextualization with LLMs improves comprehension of salient excerpts from long-form social conversations. Using GPT-4o, the authors compare raw excerpts with two contextualized versions: CEEi, generated zero-shot from the full transcript, and CEEe, generated with explicit instructions to include socially grounded attributes such as speaker characteristics, motivation, and relevant locations or events. Both contextualized versions receive significantly higher ratings than the original excerpts on textual-quality and speaker-perception dimensions (p<0.05, Welch t-test), and CEEe generally matches or beats CEEi on readability, low redundancy, and empathy while being designed to be more focused and less redundant. The paper also reports that faithfulness to the source conversation is high for short factual attributes but weaker for open-ended background and personal experiences, and that summaries from context-enriched excerpts are rated more informative and coherent than summaries of the full conversation.

Load-bearing premise

The whole result rests on treating subjective Likert ratings as a measure of comprehension, and because the enriched texts are longer and more detailed than the originals, the ratings could be rewarding extra text rather than social context specifically.

Editorial extensions

If this is right

  • Context-enriched excerpts improve reader comprehension, readability, and empathy compared with raw salient excerpts, so excerpt-sharing workflows in civic dialogue can benefit from automatic contextualization.
  • Explicit prompting for social attributes yields contexts that are rated as readable and low-redundancy, while implicit contextualization is rated more complete, indicating a trade-off between focus and breadth.
  • Faithfulness is high for short factual attributes such as name, gender, and race but low for open-ended factors such as speaker background, personal experiences, and conversation location, so downstream use must verify those details.
  • Summaries built from context-enriched excerpts are rated more informative and coherent than full-conversation summaries, suggesting contextualized excerpts can serve as a focused summarization route.
  • The released HSE dataset provides human-annotated salient excerpts and social attributes, enabling direct comparison of future contextualization methods.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether the gains come from social attributes or simply from more text; a length-matched control would settle this, and such a test is a natural next step.
  • The observation that readers can disagree with a speaker while still rating empathy and perspective-taking highly suggests contextualization may support understanding across disagreement; testing in polarized settings would be a direct extension.
  • If these results replicate, interfaces that share excerpts between groups could automatically prepend a short speaker-context block, but only after faithfulness for open-ended attributes improves.
  • The explicit prompting recipe could be reused to generate training data for smaller models that contextualize excerpts without reading the full transcript at inference time.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper addresses the problem that salient excerpts drawn from long-form, multi-party social conversations lose social context when shared with new audiences. The authors introduce the Human-annotated Salient Excerpts (HSE) dataset built on the Fora corpus, propose two LLM-based contextualization methods (CEEi, implicit zero-shot contextualization, and CEEe, explicit contextualization guided by social attributes such as speaker characteristics, motivation, and relevant locations/events), and evaluate the resulting context-enriched excerpts through subjective human ratings and objective faithfulness metrics. They report that both CEEi and CEEe receive significantly higher human ratings than the original excerpts across textual-quality and speaker-perception dimensions, that explicit contextualization is generally preferred, and that context-enriched excerpts yield conversation summaries rated more informative and coherent than full-conversation or excerpt-only summaries. The paper also reports that LLMs struggle with long-response social factors and releases the HSE dataset.

Significance. If the central claim is supported, the paper would make a useful contribution: the HSE dataset is a new community resource, the effective-contextualization framing is novel for long-form social conversations, and the human evaluation is substantial in scale (75 evaluators for the main comparison, 30 for the LLM comparison, 34 for summarization). The qualitative analysis and the extrinsic-knowledge probe are also thoughtful. However, the current evaluation does not isolate the mechanism claimed in the title and abstract. The subjective ratings are confounded with text length, the reported lengths for CEEe are internally inconsistent, and the objective faithfulness measure is both dependent on the same model used for generation and shows low scores on the very long-response social factors emphasized in the paper. These issues are load-bearing because the central claim is that social contextualization, not merely additional text, improves comprehension.

major comments (3)
  1. [Section 5.1 / Section 6.1, Table 1] The central comparison in Table 1 confounds contextualization with length. Original excerpts average 128 words, CEEi contexts average 267 words, and CEEe contexts for GPT-4o are reported as 296 words later in Section 6.1. Higher Likert ratings for understandability, completeness, informativeness, and detail could be produced by the added text alone, regardless of whether that text contains socially relevant attributes. The paper reports no length-matched control (for example, original excerpts augmented with non-social filler of comparable length), no factual comprehension questions, and no order-effect analysis. Because the specific claim is that social contextualization, not additional words, drives the improvement, this confound is load-bearing and must be addressed.
  2. [Section 6.1, LLM comparison paragraph] The reporting of CEEe length is internally inconsistent. The first paragraph of Section 6.1 states that CEEe contexts 'average around 55 words,' while the same section, when comparing GPT-4o, Claude Opus, and Llama 3.1-70b under the CEEe condition, reports that 'GPT-4o produced contexts averaging 296 words.' These two statements refer to the same method and the same model, so they cannot both be correct. The manuscript also concludes that 'explicit contextualization results in shorter enriched contexts' based on the 55-word figure, a conclusion that is contradicted by the 296-word figure. The reported lengths must be reconciled, and the analyses that depend on them must be rerun or reinterpreted.
  3. [Section 6.4, Table 3] The objective faithfulness evaluation is not an adequate basis for the paper's objective-comprehension claims. First, the extraction of attributes from the generated contexts is performed with GPT-4o, the same model used to generate the contexts in the main CEEe/CEEi comparison, so the metric is not an independent check on generation quality. Second, the faithfulness F1 scores for the long-response social factors that motivate the paper are low: for CEEe, location of conversation scores 0.34, personal experiences 0.48, and additional speaker background 0.69. These scores undercut the claim that LLMs successfully capture key social aspects and are not measures of reader comprehension. The paper should add an independent evaluation, such as human factual question answering about the source conversation and the generated context, and report precision and recall separately.
minor comments (6)
  1. [Section 6.2] The participant quote 'helped me empathize with the reader more' appears to be a typo for 'empathize with the speaker more'; please verify and correct.
  2. [Table 2] Table 2 uses a significance threshold of p < 0.1 for the dagger marker, whereas Table 1 uses p < 0.05; clarify whether all tests are two-sided and whether any multiple-comparison corrections were applied across the ten rating dimensions.
  3. [Table 5 / Appendix A] The header 'GP T' in Table 5 should read 'GPT'; also, the table's F1 scores are presented without precision/recall breakdowns, which would be informative.
  4. [Abstract and Section 6.4] The abstract's phrase 'objective evaluations' is misleading: the objective measures are faithfulness/consistency with the source conversation, not objective measures of reader comprehension. Please rephrase to avoid overclaiming.
  5. [Section 7, Table 4] The summarization comparison reports average ratings but no statistical significance tests, confidence intervals, or summary lengths; since the paper claims the CEE-based summary outperforms alternatives, significance testing should be reported.
  6. [Appendix D.4] The faithfulness extraction questions are numbered starting at 4; renumber them or explain the numbering convention.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's claims are empirical and do not reduce to their inputs or to self-citation.

full rationale

The paper contains no fitted parameters, no derivation chain, and no uniqueness theorem imported from the authors' prior work. The central claim—that LLM-generated contextualized excerpts receive higher human comprehension ratings—is an empirical result evaluated by crowdsourced raters against original excerpts, and the summarization result is likewise a direct comparison of three summary conditions. The two self-citations (Fora dataset by Schroeder, Roy, and Kabbara; Jiang et al. 2024 with overlapping authors) provide the dataset and related-work context, but neither is used to justify the paper's conclusions; the HSE dataset is newly annotated and the evaluations are external human judgments. The dual use of GPT-4o as both generator of contexts and extractor of attributes for the faithfulness analysis is a potential bias in an objective metric, and the reported context lengths are internally inconsistent (Section 6.1 reports CEEe averaging 'around 55 words' but later reports GPT-4o CEEe averaging 296 words), which threatens the validity of the length confound interpretation. However, these are experimental-design concerns, not cases where a prediction is equivalent to an input by construction. No circular step can be exhibited, so the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on several unvalidated assumptions: subjective ratings proxy comprehension, the 90-excerpt subset is representative, GPT-4o-based extraction is a valid faithfulness measure, and the Fora transcripts are accurate. No numeric free parameters or invented entities are introduced. The 'objective evaluation' in the abstract is weaker than the subjective evaluation because it uses a self-evaluation pipeline.

assumptions (4)
  • domain assumption Likert-scale subjective ratings of understandability, readability, and empathy are valid proxies for comprehension.
    The central claim of improved comprehension is based on these ratings; no objective comprehension questions are administered (Sections 5.1 and 6.1).
  • domain assumption The 90 annotated excerpts used in the main evaluation are representative of the HSE and Fora conversations.
    Only 90 excerpts from 35 conversations were annotated with social attributes due to cost, and selection criteria are not reported (Sections 3 and 9).
  • domain assumption GPT-4o attribute extraction with SBERT cosine similarity is a valid faithfulness measure for open-ended social attributes.
    The paper uses this for objective faithfulness, but notes low F1 for open-ended factors and that other methods were unsuitable (Sections 6.4 and 9).
  • domain assumption The Fora transcripts accurately preserve the speakers' words and social attributes.
    All contextualization and evaluation rely on transcripts rather than original audio or video (Section 3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts." pith.science (2026). https://pith.science/paper/CZ5WNLBT

@misc{pith2026241219966,
  author       = {Pith},
  title        = {Pith review of: Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CZ5WNLBT}},
  note         = {Machine review of arXiv:2412.19966}
}
read the original abstract

We focus on enhancing comprehension in small-group recorded conversations, which serve as a medium to bring people together and provide a space for sharing personal stories and experiences on crucial social matters. One way to parse and convey information from these conversations is by sharing highlighted excerpts in subsequent conversations. This can help promote a collective understanding of relevant issues, by highlighting perspectives and experiences to other groups of people who might otherwise be unfamiliar with and thus unable to relate to these experiences. The primary challenge that arises then is that excerpts taken from one conversation and shared in another setting might be missing crucial context or key elements that were previously introduced in the original conversation. This problem is exacerbated when conversations become lengthier and richer in themes and shared experiences. To address this, we explore how Large Language Models (LLMs) can enrich these excerpts by providing socially relevant context. We present approaches for effective contextualization to improve comprehension, readability, and empathy. We show significant improvements in understanding, as assessed through subjective and objective evaluations. While LLMs can offer valuable context, they struggle with capturing key social aspects. We release the Human-annotated Salient Excerpts (HSE) dataset to support future work. Additionally, we show how context-enriched excerpts can provide more focused and comprehensive conversation summaries.

Figures

Figures reproduced from arXiv: 2412.19966 by the authors.

Figure 1
Figure 1. Overview of Effective Contextualization of Salient Excerpts in Social Conversations: The figure illustrates the challenges posed when excerpts from longer social group conversations are shared without sufficient context. To address this, a large language model (LLM) is employed to generate context that incorporates key social attributes, aiming to enhance the comprehension of the excerpt, a process we term effective… view at source ↗
Figure 2
Figure 2. Number of defined terms in each category [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. a) Length of conversation (in words) in Fora dataset and b) Length of excerpts (in words) in human [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (45 more)
Figure 4
Figure 4. Figure 4: GPT Prompts Comparison Survey Demographics [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: LLM Contexts Comparison Survey Demographics [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Summary Survey 2 Demographics [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Crowd Source Survey Demographics F Screenshots of surveys F.1 Comparison of Explicit and Implicit Contextualisation Survey [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Understandability Question [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: Readability Question [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]
Figure 10
Figure 10. Figure 10: Redundancy Question [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: Completeness Question [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]
Figure 12
Figure 12. Figure 12: Cohesiveness Question [PITH_FULL_IMAGE:figures/full_fig_p023_12.png]
Figure 13
Figure 13. Figure 13: Perspective Question [PITH_FULL_IMAGE:figures/full_fig_p023_13.png]
Figure 14
Figure 14. Figure 14: Honest and Trustworthy Question [PITH_FULL_IMAGE:figures/full_fig_p023_14.png]
Figure 15
Figure 15. Figure 15: Respect Question [PITH_FULL_IMAGE:figures/full_fig_p024_15.png]
Figure 16
Figure 16. Figure 16: Empathy Question [PITH_FULL_IMAGE:figures/full_fig_p024_16.png]
Figure 17
Figure 17. Figure 17: POV Question [PITH_FULL_IMAGE:figures/full_fig_p024_17.png]
Figure 18
Figure 18. Figure 18: Missing Info Question [PITH_FULL_IMAGE:figures/full_fig_p024_18.png]
Figure 19
Figure 19. Figure 19: Clarify Terms Question [PITH_FULL_IMAGE:figures/full_fig_p025_19.png]
Figure 20
Figure 20. Figure 20: Comments Question [PITH_FULL_IMAGE:figures/full_fig_p025_20.png]
Figure 21
Figure 21. Figure 21: ICL Explanation of Terms Question [PITH_FULL_IMAGE:figures/full_fig_p025_21.png]
Figure 22
Figure 22. Figure 22: Rank Contexts Question [PITH_FULL_IMAGE:figures/full_fig_p025_22.png]
Figure 23
Figure 23. Figure 23: Justify Rank Question F.2 Crowd-Sourced Annotations for HSE [PITH_FULL_IMAGE:figures/full_fig_p026_23.png]
Figure 24
Figure 24. Figure 24: Understandability Question [PITH_FULL_IMAGE:figures/full_fig_p026_24.png]
Figure 25
Figure 25. Figure 25: Enhance Understandability Question [PITH_FULL_IMAGE:figures/full_fig_p026_25.png]
Figure 26
Figure 26. Figure 26: Further Clarification Question [PITH_FULL_IMAGE:figures/full_fig_p026_26.png]
Figure 27
Figure 27. Figure 27: Speaker Race Question [PITH_FULL_IMAGE:figures/full_fig_p027_27.png]
Figure 28
Figure 28. Figure 28: Speaker Race Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p027_28.png]
Figure 29
Figure 29. Figure 29: Speaker Occupation Question [PITH_FULL_IMAGE:figures/full_fig_p027_29.png]
Figure 30
Figure 30. Figure 30: Speaker Occupation Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p027_30.png]
Figure 31
Figure 31. Figure 31: Speaker Name Question [PITH_FULL_IMAGE:figures/full_fig_p028_31.png]
Figure 32
Figure 32. Figure 32: Speaker Name Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p028_32.png]
Figure 33
Figure 33. Figure 33: Gender Question [PITH_FULL_IMAGE:figures/full_fig_p028_33.png]
Figure 34
Figure 34. Figure 34: Speaker Gender Select Question [PITH_FULL_IMAGE:figures/full_fig_p029_34.png]
Figure 35
Figure 35. Figure 35: Speaker Gender Select 2 Question [PITH_FULL_IMAGE:figures/full_fig_p029_35.png]
Figure 36
Figure 36. Figure 36: Speaker Education Question [PITH_FULL_IMAGE:figures/full_fig_p029_36.png]
Figure 37
Figure 37. Figure 37: Speaker Education Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p030_37.png]
Figure 38
Figure 38. Figure 38: Speaker Economic Status Question [PITH_FULL_IMAGE:figures/full_fig_p030_38.png]
Figure 39
Figure 39. Figure 39: Speaker Economic Status Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p030_39.png]
Figure 40
Figure 40. Figure 40: Speaker Age Question [PITH_FULL_IMAGE:figures/full_fig_p030_40.png]
Figure 41
Figure 41. Figure 41: Speaker Age Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p031_41.png]
Figure 42
Figure 42. Figure 42: Personal Experience Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p031_42.png]
Figure 43
Figure 43. Figure 43: Motivation Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p031_43.png]
Figure 44
Figure 44. Figure 44: Location Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p032_44.png]
Figure 45
Figure 45. Figure 45: Events Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p032_45.png]
Figure 46
Figure 46. Figure 46: Context Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p032_46.png]
Figure 47
Figure 47. Figure 47: Additional Info Text Entry Question [PITH_FULL_IMAGE:figures/full_fig_p032_47.png]
Figure 48
Figure 48. Figure 48: Consent Form [PITH_FULL_IMAGE:figures/full_fig_p033_48.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references · 21 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  4. [4]

    https://support.anthropic.com/en/articles/7996885-how-do-you-use-personal-data-in-model-training How do you use personal data in model training?

    Anthropic. https://support.anthropic.com/en/articles/7996885-how-do-you-use-personal-data-in-model-training How do you use personal data in model training?

  5. [5]

    Anthropic. 2024. https://www.anthropic.com/news/claude-3-family Introducing the next generation of claude

  6. [6]

    David A Broniatowski. 2012. Extracting social values and group identities from social media text data. In 2012 IEEE 14th International Workshop on Multimedia Signal Processing (MMSP), pages 232--237. IEEE

  7. [7]

    Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint ArXiv:2005.14165

  8. [8]

    Youngjin Chae and Thomas Davidson. 2023. Large language models for text classification: From zero-shot learning to fine-tuning. Open Science Foundation

Show all 37 references
  1. [9]

    Jiaao Chen, Mohan Dodda, and Diyi Yang. 2023. https://doi.org/10.18653/v1/2023.findings-acl.584 Human-in-the-loop abstractive dialogue summarization . In Findings of the Association for Computational Linguistics: ACL 2023, pages 9176--9190, Toronto, Canada. Association for Com...

  2. [10]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  3. [11]

    Ritam Dutt, Zhen Wu, Kelly Shi, Divyanshu Sheth, Prakhar Gupta, and Carolyn Penstein Rose. 2024. https://arxiv.org/abs/2406.19545 Leveraging machine-generated rationales to facilitate social meaning detection in conversations . Preprint, arXiv:2406.19545

  4. [12]

    Guy Feigenblat, Chulaka Gunasekara, Benjamin Sznajder, Sachindra Joshi, David Konopnicki, and Ranit Aharonov. 2021. Tweetsumm--a dialog summarization dataset for customer service. arXiv preprint arXiv:2111.11894

  5. [13]

    James S Fishkin. 1997. The voice of the people: Public opinion and democracy. Yale university press

  6. [14]

    Lucie Flek. 2020. Returning the n to nlp: Towards contextually personalized classification models. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 7828--7838

  7. [15]

    Tanya Goyal and Greg Durrett. 2021. Annotating and modeling fine-grained factuality in summarization. arXiv preprint arXiv:2104.04302

  8. [16]

    Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News summarization and evaluation in the era of gpt-3. arXiv preprint arXiv:2209.12356

  9. [17]

    Herbert Paul Grice. 1975. Logic and conversation. Syntax and semantics, 3:43--58

  10. [18]

    Michael Alexander Kirkwood Halliday and Christian MIM Matthiessen. 2013. Halliday's introduction to functional grammar. Routledge

  11. [19]

    Dirk Hovy and Diyi Yang. 2021. The importance of modeling social factors of language: Theory and practice. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human language technologies, pages 588--602

  12. [20]

    Hang Jiang, Xiajie Zhang, Robert Mahari, Daniel Kessler, Eric Ma, Tal August, Irene Li, Alex Pentland, Yoon Kim, Deb Roy, and Jad Kabbara. 2024. https://doi.org/10.18653/v1/2024.acl-long.388 Leveraging large language models for learning complex legal concepts through storytell...

  13. [21]

    Dan Jurafsky, Rajesh Ranganath, and Dan McFarland. 2009. Extracting social meaning: Identifying interactional style in spoken conversation. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computatio...

  14. [22]

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199--22213

  15. [23]

    Md Tahmid Rahman Laskar, Xue-Yong Fu, Cheng Chen, and Shashi Bhushan Tn. 2023. Building real-world meeting summarization systems using large language models: A practical perspective. arXiv preprint arXiv:2310.19233

  16. [24]

    Wei Li, Wenhao Wu, Moye Chen, Jiachen Liu, Xinyan Xiao, and Hua Wu. 2022. Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods. arXiv preprint arXiv:2203.05227

  17. [25]

    Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896

  18. [26]

    Anna Mauranen. 2006. Signaling and preventing misunderstanding in english as lingua franca communication. International Journal of the Sociology of Language, 2006(177):123--150

  19. [27]

    https://platform.openai.com/docs/models/how-we-use-your-data How we use your data

    OpenAI. https://platform.openai.com/docs/models/how-we-use-your-data How we use your data

  20. [28]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084

  21. [29]

    Julian Risch, Timo M \"o ller, Julian Gutsch, and Malte Pietsch. 2021. Semantic answer similarity for evaluating question answering models. arXiv preprint arXiv:2108.06130

  22. [30]

    Deb Roy. 2023. https://www.theatlantic.com/technology/archive/2023/10/social-media-platforms-business-models-dialogue-network/675655/ The internet could be so good. really. The Atlantic

  23. [31]

    David M Ryfe. 2006. Narrative and deliberation in small group forums. Journal of Applied Communication Research, 34(1):72--93

  24. [32]

    Harold Saunders. 1999. A public peace process: Sustained dialogue to transform racial and ethnic conflicts. Springer

  25. [33]

    Hope Schroeder, Deb Roy, and Jad Kabbara. 2024. https://aclanthology.org/2024.acl-long.754 Fora: A corpus and framework for the study of facilitated dialogue . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p...

  26. [34]

    Shaochen Xu, Zihao Wu, Huaqin Zhao, Peng Shu, Zhengliang Liu, Wenxiong Liao, Sheng Li, Andrea Sikora, Tianming Liu, and Xiang Li. 2024. Reasoning before comparison: Llm-enhanced semantic similarity metrics for domain specialized text analysis. arXiv preprint arXiv:2402.11398

  27. [35]

    Diyi Yang, Dirk Hovy, David Jurgens, and Barbara Plank. 2024. The call for socially aware language technologies. arXiv preprint arXiv:2405.02411

  28. [36]

    Selvaraj

    Zonghai Yao, Benjamin J Schloss, and Sai P. Selvaraj. 2023. https://arxiv.org/abs/2310.05857 Improving summarization with human edits . Preprint, arXiv:2310.05857

  29. [37]

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can large language models transform computational social science? Computational Linguistics, 50(1):237--291

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.