An end-to-end data science project built on 4 years of US retail sales data (Superstore dataset).
- Time series decomposition and stationarity testing (ADF)
- 4 forecasting models: SARIMA, Prophet, XGBoost, LSTM
- Anomaly detection using Isolation Forest and Z-Score methods
- Product demand segmentation using K-Means clustering + PCA
- Interactive Streamlit dashboard with 5 pages including an AI-powered Q&A page
Python, Pandas, Statsmodels, Prophet, XGBoost, TensorFlow, Scikit-learn, Plotly, Streamlit
Superstore Sales Dataset (2015–2018) — 9,800 rows, 18 features