If we determine which data is actually necessary, we can build better user trust while maintaining the efficacy of a chatbot.
As the demand for mental healthcare increases, more people have turned to mental health chatbots and other AI-powered teletherapy options. While these chatbot applications expand access to care that is often expensive and stigmatized, they also pose some cybersecurity risks. Some chatbot applications use phone sensors and other device data to make predictions about the severity of a patient’s mental illness. While this makes these predictions more accurate, they might compromise user trust. In this paper, we are trying to find which data we should collect from the user’s phone to minimize the amount of data being collected but also collecting adequate information to predict their mental health score. This is important because users are not fully trusting the mobile apps since they are afraid the apps might track all of their personal information. If we determine which data is actually necessary, we can build better user trust while maintaining the efficacy of a chatbot. The overall approach was to first collect the data that was deemed useful, such as education, sensing, survey, and Ecological Momentary Assessment (EMA) data, and then found how the accuracy of the scores would decrease when each category of data collected was removed to determine the most important category to collect, as well as what combination of categories yielded the highest accuracy. The most significant result was that when we didn’t include the data from the surveys, the accuracy went down significantly. The major conclusions are that not every detail from a user’s phone will need to be collected in order to yield accurate results, and in fact, simply asking users about their daily experiences allows for far more accurate results than on-device measures.
Related Projects