Abstract
Web scraping refers to the collection of personal data and publicly available data on social media platforms for purposes like training, researching and testing products and services. Web scraping can be both legitimate when working in compliance with data protection safeguards and unethical when such data is breached for misuse of personal identifiable information. The paper analyses the web scraping policies of four major social media platforms, X, Meta, Reddit and LinkedIn to understand the extent of web scraping through third party applications and advertisers, the amount of data scraped, the purposes it is used for and the duration of such retention. The practice of web-scraping is prevalent among these intermediaries to further develop their platforms. These datasets range from private conversations and saved drafts to browser activities. Further, the paper compares the Data Protection Safeguards of four countries, India, California, the European Union and the Peoples’ Republic of China to further study the efficiency of such safeguards in regulating the policies of Social Media Intermediaries within their jurisdictions. The paper delves into the differences in application of such policies of the Big Tech and the provisions regarding right to privacy, erasure and right to be forgotten. It also provides an insight into the web scraping practices of the social media intermediaries to train their Artificial Intelligence Platforms through collection and retention of users’ data, which can be both personal and identifiable. The methodology is empirical analysis of specific policies in varied jurisdictions. The paper aims to answer whether the Digital Personal Data Protection Act, 2023 is compatible with international frameworks that address data scraping. The legitimacy of these practices varies according to different safeguards, and stricter protocols like the GDPR and CCPA ensure better protection of the user data.