Pith. sign in

REVIEW 2 cited by

Large Language Models in Analyzing Crash Narratives -- A Comparative Study of ChatGPT, BARD and GPT-4

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.13563 v1 pith:UAUYDVFK submitted 2023-08-25 cs.CL cs.AIcs.IRcs.LG

classification cs.CLcs.AIcs.IRcs.LG
keywords crashwerenarrativesinformationquestionsextractingmodelsrelated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In traffic safety research, extracting information from crash narratives using text analysis is a common practice. With recent advancements of large language models (LLM), it would be useful to know how the popular LLM interfaces perform in classifying or extracting information from crash narratives. To explore this, our study has used the three most popular publicly available LLM interfaces- ChatGPT, BARD and GPT4. This study investigated their usefulness and boundaries in extracting information and answering queries related to accidents from 100 crash narratives from Iowa and Kansas. During the investigation, their capabilities and limitations were assessed and their responses to the queries were compared. Five questions were asked related to the narratives: 1) Who is at-fault? 2) What is the manner of collision? 3) Has the crash occurred in a work-zone? 4) Did the crash involve pedestrians? and 5) What are the sequence of harmful events in the crash? For questions 1 through 4, the overall similarity among the LLMs were 70%, 35%, 96% and 89%, respectively. The similarities were higher while answering direct questions requiring binary responses and significantly lower for complex questions. To compare the responses to question 5, network diagram and centrality measures were analyzed. The network diagram from the three LLMs were not always similar although they sometimes have the same influencing events with high in-degree, out-degree and betweenness centrality. This study suggests using multiple models to extract viable information from narratives. Also, caution must be practiced while using these interfaces to obtain crucial safety related information.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges

    cs.AI 2026-08 conditional novelty 6.0 of 10

    For large multimodal agents in transportation, the evidence supports bounded orchestration of specialist tools, not replacement; no reviewed family demonstrates sustained real-world deployment.

  2. TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A prompt-only LLM framework that generates and executes Python code answers 90 to 93 percent of 100 GTFS transit-data queries without fine-tuning.

Pith tools