Back To Projects

Building an Optimized algorithm that provides summaries of legal documents

Aman B.

The legal industry is built around documents as they provide evidence and reduce doubt in the court. Due to the large volume of documentation in the legal industry, the processing and summarization of these documents is important to a number of individuals. We were able to create a user interface that allows for the input of documents and makes use of the algorithm we created to output a summary of the document which can be copied by the user.


We analyze the accuracy of various NLP algorithms in providing text summarization and fine-tune a particular model on a dataset to provide accurate text summaries of legal documents. The legal industry is built around documents as they provide evidence and reduce doubt in the court. Due to the large volume of documentation in the legal industry, the processing and summarization of these documents is important to a number of individuals. For example, lawyers, clients and professionals may need access to summaries of the documents for reference to similar cases. In this paper, we have developed an algorithm that trains the T5 model on the legal domain in order to create more accurate summaries of legal documents. We were able to create a user interface that allows for the input of documents and makes use of the algorithm we created to output a summary of the document which can be copied by the user.

Explore More!

Source Code
Aman B.
Eric Bradford
Electrical Engineering and Computer Science Masters from MIT, Technical PM at Apple

Related Projects

Can A Person’s MBTI Type Be Determined By A Sample Of Their Writing?

This project aims to lessen the reliance of self-report for personality typology through an artificial intelligence algorithm that can type people as one of the 16 MBTI types using an unedited writing sample by that person.
Parinita K.
Mentored by Philip Bell
Applying Machine Learning to Historical Cyber Attack Data to Predict and Prevent Future Attacks

This research demonstrates that machine learning algorithms can leverage historical data to predict the nature and severity of future cyber attacks with reasonable accuracy, revealing that despite the sophistication of state-sponsored actors, their attack patterns often exhibit digital fingerprints that offer valuable clues for anticipating future threats amidst global geopolitical tensions.
Isabella F.
Mentored by Simeon Sayer
Exposing Undercounts in the Census through Regression Modelling

Although many community leaders have proposed that language barriers pose significant obstacles to Census outreach, this paper explores the viability of using predictive models to quantify the extent the role language plays.
Tarun S.
Mentored by Katie O'Nell