Back To Projects

Predicting Future Phonological Changes Of Mandarin Chinese

Peijie G.

We created a system of translating IPA representation to vectors that captures characteristics of phonemes such as their articulatory location, and experimented with several machine learning models to capture existing trends from Old, Middle to Modern Chinese (Mandarin).


Linguists made historical reconstructions extensively based on existing patterns and common trends of phonological changes without the aid of sound data, yet previously little was done to extend such patterns into the future as a result of mostly human work in the field of historical linguistics. We created a system of translating IPA representation to vectors that captures characteristics of phonemes such as their articulatory location, and experimented with several machine learning models to capture existing trends from Old, Middle to Modern Chinese (Mandarin). Then we attempted to predict future Mandarin pronunciations based on existing phonological data and ancient reconstructions. The models were successful at capturing trends of Ancient to Modern Chinese shifts, as a great majority of the validation results were exactly correct. The predictions from the models vary in consistency of outputs and phonological inventory, but they are largely logical and may offer insight into future shifts. Therefore it is possible to predict phonological changes with machine learning, yet the accuracy of the predictions still needs validation in the future. If the predictions are proven to be effective in the future, the solutions may provide a method of capturing phonological changes and validating previous ancient reconstructions.

Explore More!

Peijie G.
Kush Khosla
MS Computer Science from Stanford, Lead Data Scientist at Retain.ai

Related Projects

Is GPT-3 smarter than a sixth-grader?

Question answering (QA) and Large Language models (LLM) have been a major research focus in Artificial Intelligence for several years. In 2017, a task called Textbook Question Answering (TQA) was introduced. The task included lessons from a middle school science textbook consisting of texts, diagrams, and natural questions. Many people attempted to create question answer models but reported sub-par accuracies.
Anitej S.
Mentored by Eric Bradford
Generating Instagram Captions with ViT-GPT2 and GPT3

This paper presents a novel approach for generating Instagram captions based on visual features and language models. Our caption generator combines Vision Transformer-GPT2 and GPT3 to generate descriptive and engaging captions in the style of an Instagram post.
Ariel M.
Mentored by Roger Jin
Can A Person’s MBTI Type Be Determined By A Sample Of Their Writing?

This project aims to lessen the reliance of self-report for personality typology through an artificial intelligence algorithm that can type people as one of the 16 MBTI types using an unedited writing sample by that person.
Parinita K.
Mentored by Philip Bell