INVESTIGATION | PRIVACY & INFLUENCE
An investigation into what major social platforms collect, what their automated systems infer and how those conclusions shape what people see.
By Andrew McDonald · 13 July 2026 · 12 min read · Policy references rechecked 28 July 2026
The information you give is only the beginning
“The government has all the information I am ever going to give them.”
I overheard that sentence in an ordinary conversation. It sounded settled, almost reassuring. The speaker seemed to imagine personal data as a finite collection of facts: a name, address, tax number, licence, medical record and perhaps a few forms completed over a lifetime.
But the most revealing information about a person is no longer limited to what they deliberately hand over.
A social platform can observe what someone watches, what they skip, which search they repeat, whose profile they revisit, where their device appears to be, what they buy elsewhere and how their behaviour changes over time. Automated systems can combine those signals and produce conclusions the person never typed into a profile.
The short answer is that platforms collect data through their services, devices, partners and tracking technologies. AI then helps turn those records into predictions about identity, interests, age, preferences and likely behaviour.
That distinction matters. AI is not a separate creature quietly vacuuming information from a phone. Companies collect data through products and commercial networks. Recommendation systems, advertising systems, analytics and generative AI then make that data more useful, more scalable and, in some cases, more intimate.
Three versions of you
European data-protection guidance offers a useful way to understand the process. It separates social-media data into three broad categories: provided, observed and inferred.
1. The person you describe
Provided data is what you actively submit: your name, age, employer, photographs, posts, comments, contacts and anything you choose to tell an AI feature. Even here, the information may describe other people who never agreed to be part of the record.
2. The person your behaviour reveals
Observed data is created through use: videos watched, links opened, pauses, searches, ad interactions, location signals, device identifiers and activity on other websites or apps that use a platform’s advertising technology.
3. The person the system predicts
Inferred data is created from the first two layers. A system may assign interests, demographic ranges, likely purchasing intentions or other classifications. The European Data Protection Board says social-media targeting may use provided, observed or inferred data, or a combination of all three.
The ability to infer sensitive traits is not new. In 2013, researchers showed that Facebook Likes from more than 58,000 volunteers could predict a range of private attributes, including political views, religious views and personality traits. The study does not prove that today’s platforms make every one of those predictions about every user. It demonstrates the underlying point: apparently ordinary digital traces can reveal more than the person intended.
What six major platforms say they collect and infer
The following comparison is based on public platform policies originally reviewed on 13 July 2026 and rechecked for material changes before publication on 28 July 2026. Wording and controls can differ by country, age, account type, product and setting.
| Platform | Examples of data observed | Stated inferences or AI-related uses |
|---|---|---|
| Meta, Facebook and Instagram | Posts, follows, engagement, watch activity, partner and advertising-tool data | Content and ad recommendations; AI training from public adult content and AI interactions, subject to region and controls |
| Google and YouTube | Searches, videos watched, content and ad interactions, device and location data, activity on partner sites and apps | Recommendations, personalised services and ads, automated content analysis and pattern recognition |
| TikTok | Watch and search activity, content during creation, image and audio features, device details and partner activity | Interests and demographic inferences in some regions; content and ad recommendations; machine-learning improvement |
| X | Posts, views, listens, Direct Messages, device and log data, partner activity and some signed-out activity | Inferred identity, recommendations and ads; training of machine-learning and AI models |
| Profile and career data, searches, content read, job activity, messages where settings allow and partner data | Industry, seniority, compensation bracket, age, gender and interests; AI model training and insights | |
| Snapchat | Content and metadata, Stories watched, Memories, My AI interactions, device sensors, location and advertiser data | Interest inference, content and ad personalisation, machine-learning development and My AI improvement |
Meta: an AI conversation can become a recommendation signal
Facebook and Instagram have long learned from visible behaviour such as follows, likes, comments, watch activity and engagement. Meta also receives information from partners and from businesses using its advertising tools. The company uses those signals to rank content, recommend accounts and personalise advertising, subject to region and settings.
Generative AI adds another layer. Meta says it uses public posts and comments shared by adults, together with interactions people have with Meta AI, to train and improve its AI models. In the European Union, adults can object to the use of their public content for this training. Meta says private messages with friends and family are not used for AI training unless someone chooses to share those messages with an AI feature.
Training is not the only use. From December 2025, Meta began using text and voice interactions with its AI features as signals for content and advertising recommendations in most regions. Its own example is simple: ask Meta AI about hiking and you may later see hiking groups, trail posts or advertisements for boots. Meta says it does not use certain sensitive topics from AI conversations to show ads.
The important boundary is therefore not simply public versus private. A conversation may feel personal while still becoming product data, a training input or a recommendation signal, depending on what was shared, where it occurred and which regional rules apply.
Google and YouTube: activity can travel across services
Google’s privacy policy says it may collect search terms, videos watched, interactions with content and ads, purchases, people with whom someone communicates, activity on third-party sites and apps that use Google services and synced Chrome history. Depending on settings, it can combine information across services and devices.
The policy also says automated systems analyse content to provide tailored search results, personalised ads and other features, and that algorithms recognise patterns in data. YouTube watch and search history can shape recommendations. Activity on a non-Google site or app may be associated with personal information when relevant account controls allow it.
There are stated limits. Google says it does not show personalised ads based on sensitive categories such as race, religion, sexual orientation or health, and does not personalise ads from the contents of Drive, Gmail or Photos. Those protections matter, but they do not make the broader activity record disappear. Search, watch behaviour and service interactions can still create a detailed picture of attention and intent.
TikTok: the draft, the face and the pause
TikTok publishes different privacy policies for different regions, so the exact permitted uses and controls depend on where a person is located. Its policies describe extensive collection of viewing, search and browsing activity, content, device details, location information where permitted and information from other sources.
Some regional policies also describe automated analysis of images, video and audio and the use of data to personalise content and advertising and improve machine-learning systems. The specific wording and categories vary by jurisdiction, which is why a single global description can be misleading.
The broader point remains: the permitted data environment can extend well beyond the finished video someone believes they chose to share.
X: public content is only one part of the record
X describes itself as a public platform, but its policy covers more than public posts. It lists viewing and listening history, likes, bookmarks, downloads, follows, Direct Message contents and metadata, device information, approximate location, advertiser data and information about activity on partner websites and apps.
X also says it may receive log information when someone views or interacts with its services without an account or while signed out. That can include IP address, pages visited, device and application identifiers, ads shown and search terms. The policy says X may infer identity by associating devices, browsers and identifiers.
The AI connection is explicit. X says it may use collected and publicly available information to help train machine-learning or AI models.
LinkedIn: a professional profile becomes a prediction
LinkedIn begins with information people expect to be professional: employment history, education, skills, applications, connections and activity. It also collects searches, content read, pages visited, videos watched, advertising interactions and information from partners and publishers.
Its policy provides unusually concrete examples of inference. LinkedIn says it may use a job title to infer industry, seniority and compensation bracket; a graduation date to infer age; a first name or pronoun use to infer gender; feed activity to infer interests; and device information to recognise a member.
LinkedIn also says it may use personal data to develop and train AI models and to generate insights through AI, automated systems and inferences.
The result can be a second professional identity that the person did not write: a predicted level of seniority, earning band, age or interest profile. Even an inaccurate inference may shape which ads, opportunities or recruiter tools place them in view.
Snapchat: private messages have limits, but AI chats are different
Snapchat draws an important line between ordinary communications with friends and interactions with My AI. Snap says My AI conversations are retained until users delete the content or their account, and that the content shared with My AI can be used to improve Snap products and personalise the experience, including ads.
The difference between a chat with a friend and a chat with an AI may not feel large on the screen. In the policy, it can be decisive.
The tracking does not stop at the edge of the app
A person does not have to describe a purchase on social media for a platform to learn about it. Advertising pixels, software kits, cookies and partner feeds can send information from other websites, apps and stores back into advertising systems.
Australia’s privacy regulator describes a tracking pixel as code placed on a website that can send activity to a third-party provider. Depending on the implementation, that activity may include pages visited, clicks, items placed in a cart, IP address, geolocation, URL information and information entered into forms.
This is why an advertisement can appear to know about something never posted. The explanation may be less dramatic than a microphone secretly listening and more systematic: a purchase event, location signal, page view or shared identifier has been matched to a profile, which is then placed into an audience or interest category.
What a privacy policy can and cannot tell us
A privacy policy is evidence of what a company says it collects, uses or may do. It is not a live map of every database, a measurement of how often each field is used or an independent audit of whether every safeguard works as described.
The US Federal Trade Commission reached beyond public policies by compelling information from nine major social and video services. Its 2024 staff report found extensive collection about users and non-users, data from brokers, broad sharing and weak data-minimisation and retention practices. It also found that people often had little or no way to opt out of their information being fed into automated systems, while approaches to monitoring and testing those systems were inconsistent.
The report does not mean every platform behaves identically or that every use is unlawful. It does show why reading a consent screen is not the same as holding the system accountable. The public usually sees a policy and a handful of settings. The company sees the data flows, model features, experiments, error rates and commercial value.
Inference changes the privacy question
Traditional privacy advice tells people to share less. That remains useful, but it is no longer sufficient.
A person can avoid listing a political belief and still watch the same speakers repeatedly. They can withhold their age and still provide a graduation date. They can keep a concern private while searching, pausing and engaging in ways that make the concern statistically visible.
As our related investigation AI Knows What You Fear, Want and Regret examines, an inference may be correct, wrong or only weakly probable. All three can matter. A correct inference reveals something the person chose not to state. A wrong inference can still place them in the wrong audience, alter recommendations or shape how a system responds. A probability can be treated as a fact once it enters a large automated process.
People who are young, elderly, distressed, socially isolated or less confident with technology may be less able to recognise targeting or challenge an automated classification. They may also reveal more to an AI feature because it feels private, patient or helpful.
The answer cannot be to blame the person who clicked accept. Meaningful control requires limits on unnecessary collection, clearer separation between service functions and advertising, accessible explanations of important inferences, reliable deletion and independent scrutiny of automated systems.
What people can do now
Review the profile the platform shows you
Open advertising, privacy and personalisation settings. Look for inferred interests, connected accounts, off-platform activity, location history, contact uploads and permissions. The labels differ, but the categories are usually recognisable.
Download your data
Most large platforms provide an access or download tool. The export may not reveal every model feature or internal inference, but it can expose forgotten searches, contacts, devices, locations and activity histories.
Treat AI conversations as data
Before sharing sensitive information with a built-in assistant, check whether the conversation is retained, used for personalisation or used to improve models. Use temporary or non-training modes where available, and remove names or details about other people when they are not necessary.
Reduce the signals you do not need to provide
Disable precise location, contact syncing, cross-app tracking and unnecessary device permissions when they are not needed for a feature you value. Clear or pause watch and search histories if the service provides that control.
Do not mistake settings for complete control
Settings can reduce collection or personalisation, but they do not necessarily reveal or erase every inference already created. Rights to access, object, correct or delete data also vary by jurisdiction. Where a platform’s response is inadequate, a national or regional privacy regulator may provide a complaint route.
You did not have to tell them
Government records and commercial platform profiles are not the same. Governments may hold official information under legal authority. Platforms can observe everyday attention, relationships, movement and commercial behaviour at a scale that produces a different kind of knowledge.
The speaker I overheard was right about one thing: there may be facts they never intend to give anyone again.
But a system does not always need the confession. A pattern can be enough.
The same issue becomes even more personal when systems can reproduce identity, as examined in Who Owns Your Face?
The harder privacy question is no longer only, “What did you tell them?” It is who is allowed to decide what your behaviour means, how long that conclusion follows you, and what choices are quietly made for you because of it.
Editorial note: This article compares public policies and published regulatory or academic evidence. It does not claim that every listed data type is used for every person or in every jurisdiction. It is analysis, not legal advice or original technical auditing.
AI disclosure: This article was developed with assistance from artificial intelligence. Its sources, claims and conclusions were reviewed and approved by Immortal AI’s editor.
