Machine Learning and Privacy in Social Networks
This paper examines the growing tension between machine learning advancements and user privacy on social networks. Drawing on a review of recent literature, it explores how machine learning algorithms are being deployed both to harvest personal data — predicting user behavior, political affiliations, and habits — and to protect that data through privacy-preserving protocols such as Chiron, CodedPrivateML, and secure aggregation frameworks. The paper discusses the ethical implications of Big Data collection, the risk to health data privacy, and the inadequacy of existing legal protections. It concludes with recommendations for targeted legislation governing machine learning use on social networks and for broader public education on digital privacy practices.
- Introduction: Machine learning threatens social media user privacy
- Privacy and Machine Learning in Social Networks: How algorithms collect and expose personal user data
- Privacy-Preserving Protocols and Frameworks: Technical solutions designed to protect user data
- Ethical Implications and Big Data: Ethics of Big Data collection and HIPAA risks
- Findings and Recommendations: Legislation and education recommended to protect privacy
- Conclusion: Ongoing data war requires urgent legislative action
✍️ How to write this paper — guide, tools & examples ▾
What makes this paper effective
- Synthesizes a range of technical research sources into a coherent narrative accessible to a non-specialist audience, balancing depth with clarity.
- Maintains a balanced perspective by acknowledging both the beneficial and harmful applications of machine learning, avoiding one-sided advocacy.
- Grounds abstract technical concepts (e.g., secure aggregation, sandboxing) in concrete implications for ordinary social media users.
- Closes the loop between literature review and actionable recommendations, connecting research findings directly to policy and education proposals.
Key academic technique demonstrated
The paper demonstrates effective thematic synthesis of a literature review. Rather than summarizing each source in isolation, it groups studies by the function they serve — privacy protection versus data collection — and uses them to build a cumulative argument about an ongoing, unresolved conflict. The Cold War analogy for escalating machine learning arms races is a strong rhetorical move that gives the reader a conceptual framework for the stakes involved.
Structure breakdown
The paper follows a conventional research-paper structure: an abstract and introduction frame the problem; a literature review occupies the bulk of the body, organized around privacy threats and then privacy-preserving countermeasures; a findings and recommendations section synthesizes the evidence into practical proposals; and a brief conclusion reinforces the central argument about legislative urgency. The structure is logical and well-signposted throughout.
Introduction
Social networks have allowed an ocean of personal data to form that is now sitting waiting for machine learning algorithms to collect, analyze, and use to recognize individuals on social media (Oh, Benenson, Fritz & Schiele, 2016). Machine learning algorithms are thus being used more and more in social networks to collect data on users and to assess their browsing and personal information — and in doing so they could soon be predicting someone's recreational activities or political affiliation through a simple analysis of an individual's social media use, such as posts on Twitter or the friends one has on Facebook (Lindsey, 2019). As a result, the privacy of individual social media users may be in jeopardy. This paper reviews the findings of the related literature on this subject and discusses them along with recommendations for addressing this issue in the future.
Privacy and Machine Learning in Social Networks
Privacy and information sharing may seem like two diametrically opposed concepts in the context of social media, and to a high degree they are. Mobile devices allow users to set information-sharing settings that enable algorithms on other applications to identify a person's location, habits, and other information in order to personalize advertisements and so on. The information available for viewing by machine learning programs is enormous, and many users do not even realize it. Machine learning algorithms often know more about a user's habits and choices than the user does themselves.
Bilogrevic et al. (2016) point out that "by analyzing people's sharing behaviors in different contexts, it is shown in these works that it is possible to determine the features that most influence users' sharing decisions, such as the identity of the person that is requesting the information and the current location" (p. 126). The problem is that users do not know how to articulate their own personal information-sharing preferences, nor are these preferences static. For that reason, Bilogrevic et al. (2016) created a program that uses machine learning AI to automatically configure those settings based on user interaction on the web, reducing the user's uncertainty about when it is appropriate to share information and when it is not. The program developed by Bilogrevic et al. (2016) handles that decision for them.
Such a program is one example of the way privacy concerns and machine learning advancements are meeting in the realm of social media. One reason it is needed is that, as Oh et al. (2016) show, there is effectively no such thing as privacy on the Internet, and AI is being developed to collect as much data as possible on what is publicly available. The implications for personal privacy are enormous, especially as more and more people put their entire lives online (Oh et al., 2016). Yet while there are machine learning programs being developed to collect information on users — in order to recognize them, create profiles of them, and predict their behaviors — there are simultaneously machine learning programs like the one developed by Bilogrevic et al. (2016) being designed to protect and preserve users' privacy.
Privacy-Preserving Protocols and Frameworks
Mohassel and Zhang (2017) use a two-server model with machine learning algorithms for the purposes of training linear regression, logistic regression, and neural network models. In this model, during a setup phase, "the data owners (clients) process, encrypt and/or secret-share their data among two non-colluding servers. In the computation phase, the two servers can train various models on the clients' joint data without learning any information beyond the trained model" (Mohassel & Zhang, 2017, p. 19). Their protocol is 1,100–1,300 times faster than previous protocols developed to protect end users' privacy in an Internet environment where scanning algorithms are constantly seeking user data to build profiles. Their protocol acts as a wall between the user and other machine learning systems seeking their information.
Bonawitz et al. (2017) examine some of the reasons machine learning models have attracted so much interest in recent years. They note that these models have practical applications that can be used to enhance the public good — for instance, they can facilitate "everything from medical screening to disease outbreak discovery" (Bonawitz et al., 2017, p. 1175). This is important to remember because in the debate over whether machine learning models are beneficial or harmful, and what their ramifications for society might be, their potential uses will be part of that discussion; any ignorance of the possible benefits of their application risks being judged as bias. There is therefore a need to balance the positive returns this technology could produce against the potential harms it could cause. Somewhere within that discussion lies the issue of privacy, but how it fits into the nexus of technology for safety versus technology to preserve autonomy and privacy remains unclear and subject to debate.
One case in point regarding the importance of ensuring that data collection instruments are not abused or used recklessly is provided by Bonawitz et al. (2017), who describe the case of a U.S. Supreme Court nominee whose video rental history was mishandled and released in 1988 without his consent, damaging his reputation in the process. The fallout from that incident prompted new legislation, and the law "passed in response to that incident remains relevant today, limiting how online video streaming services can use their user data" (Bonawitz et al., 2017, p. 1175). However, laws are one thing and working technology is quite another — and that is where the protocol from Bonawitz et al. (2017) comes into play. Their technology is a protocol that securely computes sums of various vectors, for a constant number of rounds, with low communication overhead and robustness to failures, using "only one server with limited trust" (Bonawitz et al., 2017, p. 1176). The outcome is that protection is provided against honest-but-curious (passive) adversaries.
Hunt, Song, Shokri, Shmatikov, and Witchel (2018) designed a system that provides privacy-preserving machine learning as a service to clients storing data in the cloud. Their protocol, named Chiron, preserves privacy by hiding training data from the service operator while also hiding the algorithm and structure of the model from the end user, effectively leaving only "black-box access to the trained model" (Hunt et al., 2018, p. 1). Their method is based on the Ryoan sandbox, designed to prevent data leaks: "to enforce data confidentiality while allowing the provider to select, configure, and train a model any way they want, Chiron employs a Ryoan sandbox, which in turn is based on a hardware-protected enclave such as Intel's SGX" (Hunt et al., 2018, pp. 1–2). The model works, but the problem of advances by adversaries remains an ongoing issue.
So, Guler, Avestimehr, and Mohassel (2019) address the challenge of training a machine learning model while keeping user data private and secure. Their solution, CodedPrivateML, "keeps both the data and the model information-theoretically private, while allowing efficient parallelization of training across distributed workers" (p. 1). Similar in spirit to what Hunt et al. (2018) achieved with Chiron, So et al. (2019) accomplish with CodedPrivateML. The similarity of approach across these training models and the persistent need to maintain privacy demonstrates that this issue is far from resolved.
Balle et al. (2019) likewise focus on the problem of machine learning and privacy rights. They discuss the merits of various approaches in the realms of cryptography, security, and machine learning, with a view to identifying which methods are most efficient for preserving privacy. While helpful, the overall discussion illustrates why the issue is unlikely to be resolved anytime soon: as quickly as solutions and protections can be developed, they can be overcome, because machine learning advances are being made on both sides. It is analogous to the Cold War, where weapons stockpiling occurred because neither power could afford to pause lest a vulnerability be perceived. At some point there will be a massive gap between those who are protected and those who are highly vulnerable — having failed to keep pace with these advances.
Conclusion
Social networks are vulnerable to data harvesting algorithms, and machine learning has only advanced the degree to which personal data may be gathered and interpreted for predictive purposes. The anticipated behaviors of users can be predicted based on past inferences by machine learning algorithms, but machine learning can also be used to protect the privacy of users and share only the information that the user is comfortable sharing. Machine learning protocols have even been developed that conceal data from Big Data harvesters. And yet there is no end in sight for this war over data — the conflict between what should be public and what should be private. Until legislation is passed to settle the debate, the war is likely to escalate.
Create your account
Always verify citation format against your institution’s current style guide requirements.