Traffic Accident Severity Prediction
Route-risk prediction trained on 100,000+ accident records
- Role
- Data science & backend
- Period
- 2025
- Status
- Open source
- Links
- GitHub ↗
By the numbers
100K+100K+ - accident records
85%+85%+ - accuracy
33 - models compared
Overview
Route-risk prediction trained on 100,000+ historical accident records, issuing real-time alerts at 85%+ accuracy using OpenRouteService and weather APIs.
- Logistic regression, Random Forest and XGBoost compared; evaluated on accuracy, F1 and ROC-AUC.
- A Flask API joins route segments with live weather and returns a risk score; a React front end shows it on the map.
01
Problem
Most factors that determine accident severity (weather, hour, road type, visibility) are known before the trip starts. The question: can that information produce a route-based, preventive alert for the driver or the authority?
02
Approach
More than 100,000 historical accident records were cleaned; new features were derived from datetime, weather and location fields. Logistic regression, Random Forest and XGBoost were trained and compared on accuracy, precision, recall, F1 and ROC-AUC.
A Flask REST API splits the OpenRouteService route into segments and predicts for each one with current weather. A React front end colours the risky segments on the map.
03
Architecture
- Data
- 100K+ kaza kaydı
- Öznitelik mühendisliği
- Model
- Scikit-Learn
- XGBoost
- Değerlendirme
- Serving
- Flask API
- OpenRouteService
- Hava durumu API
- UI
- React harita
04
Key decisions
- 01
Split the route into segments
A single 'this route is risky' score is useless. Segment-level prediction shows which part is risky and why.
Outcome
The best model reached over 85% accuracy. Experiments are reproducible in Jupyter notebooks; the API and the front end live in separate folders in the repository.
Related projects
- AI / MLOpen source
Appliances Energy Prediction
Comparing regression models for smart-home energy consumption
Predicting appliance energy use from 19,735 ten-minute readings in the UCI dataset. Random Forest, XGBoost, LightGBM and linear regression were compared; LightGBM led with R² ≈ 0.76.
- 19.735
- sensor rows
- 0.76
- R² (LightGBM)
- Python
- Pandas
- Scikit-Learn
- LightGBM
- XGBoost
- Jupyter
Open - AI / MLOpen source
Mobile Price Classification
Classifying phone price tiers from hardware specifications
A classification study predicting a phone's price tier from hardware specs such as RAM, battery, screen and camera; includes a data-collection script, Jupyter analysis and a containerized prediction service.
- Python
- Pandas
- Scikit-Learn
- Jupyter
- Docker
Open
AI / MLLiveChefsStack
Recipe platform with a vector-based recommendation engine
A recipe and recommendation platform where suggestions come from vector similarity instead of hand-written rules, and user interactions flow through Kafka and Redis into personalized feeds.
- <50ms
- similarity query
- 1.000+
- interaction events processed
- Python
- FastAPI
- PostgreSQL
- pgvector
- Redis
- Kafka
Open