We created a system of translating IPA representation to vectors that captures characteristics of phonemes such as their articulatory location, and experimented with several machine learning models to capture existing trends from Old, Middle to Modern Chinese (Mandarin).
Linguists made historical reconstructions extensively based on existing patterns and common trends of phonological changes without the aid of sound data, yet previously little was done to extend such patterns into the future as a result of mostly human work in the field of historical linguistics. We created a system of translating IPA representation to vectors that captures characteristics of phonemes such as their articulatory location, and experimented with several machine learning models to capture existing trends from Old, Middle to Modern Chinese (Mandarin). Then we attempted to predict future Mandarin pronunciations based on existing phonological data and ancient reconstructions. The models were successful at capturing trends of Ancient to Modern Chinese shifts, as a great majority of the validation results were exactly correct. The predictions from the models vary in consistency of outputs and phonological inventory, but they are largely logical and may offer insight into future shifts. Therefore it is possible to predict phonological changes with machine learning, yet the accuracy of the predictions still needs validation in the future. If the predictions are proven to be effective in the future, the solutions may provide a method of capturing phonological changes and validating previous ancient reconstructions.
Related Projects