top of page
CASE STUDIES
Classwork completed across Business Analytics I & II, Programming Essentials, Decision Analytics & Data Mining
College Application Prediction Using Artificial Neural Networks
Tools: R | Neural Networks | Regression Trees | Predictive Modeling | Data Preprocessing
Project Overview
Developed machine learning models to predict college application volume using institutional characteristics such as out of state, expenditure, class rank, private vs public and degree type. This approach would be useful in HR hiring or predicting candidate job fit where they might use variables such as tenure, reviews by peers, alma mater and previous salary history in assessments. This case included multiple neural network architectures to improve predictive accuracy and identify nonlinear relationships within higher education data.
Analytics Process
-
Prepared training, validation, and testing datasets with min-max feature scaling.
-
Tuned hidden-layer architecture and training thresholds to optimize generalization without overfitting.
-
Evaluated model performance using Out-of-Sample R² (OSR²).
Key Insights
-
Neural networks captured complex nonlinear relationships that traditional regression trees could not.
-
The optimized ANN demonstrated that institutional characteristics can accurately predict future application demand, making it suitable for enrollment forecasting and resource planning.
Baseball Salary Prediction Using Regression Trees
Tools: R | Regression Trees | Linear Regression | Predictive Modeling
Project Overview
Built predictive models to estimate Major League Baseball player salaries using performance statistics and metrics such as runs, hits, at-bats, errors, home runs, division and RBIs. This setup could be transferrable to accounting where evaluating the worth of an asset is crucial when justifying an investment. After splitting the data, we tuned regression tree complexity through cross-validation and compared results against a linear regression baseline.
Analytics Process
-
Evaluated performance using OSR² and MAE.
Key Insights
-
Career performance metrics were stronger predictors of salary than single-season statistics.
-
Regression trees effectively identified salary thresholds and nonlinear decision patterns that linear regression failed to capture.
-
Tree-based models provided greater interpretability while delivering superior predictive performance.
Results
Model Performance
Regression Tree OSR² = 0.6337
Linear Regression OSR² = 0.5385
MAE = 0.4105
Regression trees improved predictive accuracy by 17.7% over linear regression.
Movie Recommendation Engine Using Association Rule Mining
Tools: R | Association Rules | Apriori Algorithm | Market Basket Analysis | Recommendation Systems
Project Overview
Developed a movie recommendation engine using the Apriori algorithm to uncover viewing patterns and generate personalized recommendations from user movie preferences. This is highly applicable to marketing where segmenting and positioning items is imperative to generating sales and meeting customer preferences. We analyzed over 665 users and 4,035 movies to identify high-confidence recommendation rules.
Analytics Process
-
Explored user viewing behavior and movie popularity.
-
Generated associationthrough support, confidence, and lift metrics.
-
Evaluated recommendation quality through recommendation scenarios and affinity analysis.
Key Insights
-
Popular movie franchises (e.g., The Godfather, Star Wars, and The Lord of the Rings) generated the strongest association rules, indicating highly predictable sequel and franchise viewing behavior.
-
High-lift rules revealed meaningful cross-movie preferences beyond simple popularity, uncovering hidden relationships between classic films such as Vertigo, Citizen Kane, and Apocalypse Now.
Results
-
Generated 440 high-confidence association rules
-
Built recommendations with up to 85% confidence
-
Identified users who liked X2: X-Men United were 3.5× more likely to also enjoy The Matrix than the average viewer.
Supply Chain Optimization for Golf Club Production
Tools: Excel Solver | Linear Programming | Supply Chain Analytics | Optimization Modeling
Project Overview
Developed an optimization model to maximize monthly profit for a golf club manufacturer by determining the optimal production and distribution plan across three manufacturing plants and three distribution centers. The model accounted for raw material constraints, holding costs, assortment product mix, and transportation costs.
Analytics Process
-
Built a linear programming optimization model in Excel Solver.
-
Incorporated production capacity and plant-specific manufacturing constraints.
-
Optimized shipping decisions based on distribution center needs.
-
Balanced demand fulfillment requirements while maximizing total profit.
-
Validated resource utilization and distribution flows through sensitivity analysis.
Key Insights
-
The optimal solution specialized production by plant, leveraging each facility's resource availability and manufacturing capabilities to maximize efficiency.
-
Low-cost shipping routes significantly reduced transportation expenses while still satisfying distributor demand requirements.
-
Resource constraints (particularly titanium, aluminum, and rock maple availability) became the primary drivers of production decisions, highlighting production bottlenecks and opportunities for capacity expansion.
-
Optimization demonstrated how integrated production and logistics planning can substantially increase profitability while maintaining customer service levels.
Results
Total Revenue = $1.56M
Total Shipping Cost = $157K
Total Profit = $1.40M
Demand Fulfillment = 90–100%
Transportation Network Optimization for Motor Distributor
Tools: Excel Solver | Linear Programming | Network Flow Optimization | Operations Research
Project Overview
Designed and optimized a transportation network for an engine manufacturer to minimize total production and shipping costs across three manufacturing plants, two warehouses, and three distribution centers. The optimization model incorporated production capacity, minimum production requirements and warehouse flow conservation.
Analytics Process
-
Formulated the problem as a minimum-cost network flow model.
-
Developed mathematical decision variables, objective function, and operational constraints.
-
Modeled production capacities, warehouse balance equations, distributor demand, and non-negativity constraints.
-
Implemented the optimization model to find optimal distribution strategy.
-
Validated network flows to ensure demand satisfaction and operational feasibility.
Key Insights
-
Warehouse flow balancing minimized unnecessary transportation by routing shipments through the most cost-efficient distribution paths.
-
Integrating manufacturing and transportation decisions reduced total operating costs compared to optimizing production or shipping independently.
-
The network model identified cost-effective shipping routes that satisfied distributor demand without exceeding warehouse processing limits.
Results
The optimized network achieved 100% demand fulfillment while minimizing total manufacturing and transportation costs to $20,150.
Care Management Decision Making for Hospitals
Risk Classification Project in Python
Aided in heart disease detection and prediction given a patient's clinical attributes (age, sex, chest pain type, resting BP, cholesterol, fasting blood sugar, resting ECG, max heart rate, exercise-induced angina and more. The real focus rigorously estimating how well each candidate algorithm would generalize to new patients, and who to prioritize first in the system based on severity and odds.
Dataset: 918 patient records, 11 features (mixed categorical/numeric), binary target `Heart Disease`.
Key Insights & Skills Demonstrated
1. Algorithm-aware preprocessing pipelines** — Used `ColumnTransformer` + `OneHotEncoder` for categorical features. Left numeric features raw for the Decision Tree (scale-invariant), but applied `StandardScaler` for Logistic Regression, since scale affects gradient-based/linear models. Each preprocessor was packaged with its estimator in a single `sklearn.Pipeline` to prevent data leakage during CV.
2. Hyperparameter tuning with GridSearchCV** — Tuned `min_samples_leaf` for the Decision Tree and `C` for Logistic Regression on a log scale, since regularization strength acts multiplicatively. Used `RepeatedStratifiedKFold` (2×10-fold) as the CV strategy for both.
3. Nested cross-validation for unbiased generalization estimates** — The centerpiece technique: an inner `GridSearchCV` selects hyperparameters per fold, while an outer `cross_validate` evaluates the tuned model on folds it never touched during tuning. This avoids the optimistic bias that comes from using the same data for both model selection and evaluation.
4. Results — comparing algorithms on equal footing**
| Model | Best Hyperparameter | Mean Accuracy (Nested CV) | Std Dev |
| Decision Tree | `min_samples_leaf = 14` | 0.836 | 0.038 |
| Logistic Regression | `C ≈ 0.089` | **0.867** | **0.027** |
Logistic Regression won on both accuracy and consistency — suggesting the relationship between these features and heart disease risk is largely linear rather than needing complex interactions.
bottom of page