Back To Projects

Predicting Repeat Purchases in E-Commerce

Ayrton S.

The underlying problem here is identifying more efficient upselling strategies for retailer companies. These findings may potentially help a retail company determine the likelihood of a customer purchasing another item based on their first purchase, and then decide how to upsell based on that information.


As companies move their business online, the e-commerce space continues to develop as a rapidly growing industry. Marketing and upselling are crucial aspects of profitability. Therefore the objective of this investigation is to determine how the price, along with other features, of items sold on Amazon affect the likelihood of a customer making a repeat purchase. The underlying problem here is identifying more efficient upselling strategies for retailer companies. These findings may potentially help a retail company determine the likelihood of a customer purchasing another item based on their first purchase, and then decide how to upsell based on that information. Customer behavior and motivation is a crucial aspect of repeat purchases and original purchases in general, this notion has to be understood before attempting to answer the question. Identifying which features held the most weight in predicting repeat purchases was also a crucial aspect in determining if customers will buy again. In general it could be deduced that price, product brand, and product rating carry a great deal of weight when it comes to customer’s making repeat purchases.

Explore More!

Ayrton S.
Bryce Johnson
Computer Science Stanford Alum, Industry Data Analytics and Business Consultant at Oliver Wyman

Related Projects

Predicting the Price of New York City Airbnbs

How can one predict the price of a New York City Airbnb? We are trying to create a machine learning model that can predict the price of a NYC Airbnb given some factors with high accuracy.
Bobby B.
Mentored by Tomer Arnon
Approaches to fraud detection on credit card transactions using artificial intelligence methods

In this paper, we study the problem of detecting fraudulent credit card transactions. We select the most relevant features using a heuristic approach, and fit three different model classes to a simulated dataset: Logistic Regression, Random Forests and Gradient Boosting Machines. We find that hyperparameter tuning has a big impact on the precision and recall of our classifiers. We also find that of the three classes, Gradient Boosting Machines were the best-performing model class, achieving 83% precision and 64% recall on unseen data.
Aryaman R.
Mentored by Yuan Lee
Predicting Fake Job Listings from Real Ones using Machine Learning Models

The main way this problem is going to be solved is through the creation of models that can predict whether these job listings are real or fake.
Ronit M.
Mentored by Hassan Azmat