Pith. sign in

REVIEW 3 cited by

Learning Traffic Crashes as Language: Datasets, Benchmarks, and What-if Causal Analyses

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10789 v1 pith:JUFEJEZU submitted 2024-06-16 cs.CV

classification cs.CV
keywords trafficcrashaccidentsenvironmentalfactorsfurtherlanguagelearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The increasing rate of road accidents worldwide results not only in significant loss of life but also imposes billions financial burdens on societies. Current research in traffic crash frequency modeling and analysis has predominantly approached the problem as classification tasks, focusing mainly on learning-based classification or ensemble learning methods. These approaches often overlook the intricate relationships among the complex infrastructure, environmental, human and contextual factors related to traffic crashes and risky situations. In contrast, we initially propose a large-scale traffic crash language dataset, named CrashEvent, summarizing 19,340 real-world crash reports and incorporating infrastructure data, environmental and traffic textual and visual information in Washington State. Leveraging this rich dataset, we further formulate the crash event feature learning as a novel text reasoning problem and further fine-tune various large language models (LLMs) to predict detailed accident outcomes, such as crash types, severity and number of injuries, based on contextual and environmental factors. The proposed model, CrashLLM, distinguishes itself from existing solutions by leveraging the inherent text reasoning capabilities of LLMs to parse and learn from complex, unstructured data, thereby enabling a more nuanced analysis of contributing factors. Our experiments results shows that our LLM-based approach not only predicts the severity of accidents but also classifies different types of accidents and predicts injury outcomes, all with averaged F1 score boosted from 34.9% to 53.8%. Furthermore, CrashLLM can provide valuable insights for numerous open-world what-if situational-awareness traffic safety analyses with learned reasoning features, which existing models cannot offer. We make our benchmark, datasets, and model public available for further exploration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A QLoRA-fine-tuned Llama3 8B model, trained on literature-derived crash factor data, generates plausible crash cause explanations for freeway crashes, validated only by 88.89% agreement from six traffic safety researchers.

  2. Predicting person-level injury severity using crash narratives: A balanced approach with roadway classification and natural language process techniques

    cs.LG 2025-09 reject novelty 4.0 of 10

    The paper claims narrative text improves person-level crash injury prediction in Kentucky, with TF-IDF and XGBoost best, but the structured-only baseline comparison is not shown.

  3. CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A fine-tuned Llama-3-8B on LLM-augmented crash narratives achieves a Macro-F1 of 0.736 for crash severity prediction, outperforming zero-shot and few-shot GPT-4o and Llama-3-70B baselines but with no error bars or cod...

Pith tools