How to run market research without LGPD risk: the complete guide to synthetic data for marketing teams

How to run market research without LGPD risk: the complete guide to synthetic data for marketing teams

Daniel Victorino

How to run market research without LGPD risk: the complete guide to synthetic data for marketing teams

Researching consumers has become riskier. The end of third-party cookies in Chrome, full LGPD enforcement, and growing ANPD oversight created an environment where any research tool that depends on identifiable personal data carries real operational risk.

The answer is not to stop researching. It is to research without touching personal data.

Why synthetic data and LGPD are being discussed at the same time

Because the end of cookies, full LGPD enforcement, and active ANPD oversight arrived together, making any research solution based on identifiable personal data a real operational risk. Synthetic data answer this pressure because they preserve behavioral intelligence without exposing real individuals.

In 2024, Google eliminated third-party cookies in Chrome, ending 25 years of large-scale individual tracking. LGPD has been in force since 2020, with ANPD enforcement intensified from 2023 onward. For marketing teams, this means consumer insight must be generated with stronger privacy protection by design.

What synthetic data are: technical and legal definition

What are synthetic data?

Synthetic data are artificially generated by algorithms that preserve the statistical properties of a real dataset without corresponding to any real individual. Legally, they are not personal data under Art. 5, I of the LGPD when they cannot identify a natural person.

The technical definition: synthetic data are created by Machine Learning models that learn statistical patterns from a real dataset and generate new data that preserve those patterns without replicating individual records.

The relevant legal definition: Art. 5, I of the LGPD defines personal data as “information related to an identified or identifiable natural person.” Synthetic data generated by Galaxies do not satisfy this condition because the output does not refer to a real individual.

Are synthetic data legal in Brazil?

Yes. Synthetic data are not personal data under Art. 5, I of the LGPD because they do not refer to an identified or identifiable natural person. No legal basis, consent, or DPA is required for the synthetic output used in market research.

The answer is yes, with an important nuance: the legality of the output (synthetic data) does not eliminate obligations over the input (real data used to train the model). Input data must have an appropriate legal basis under LGPD and be documented contractually.

In practice, for the marketing team this means consumer research available 24 hours a day, without a new collection-approval process, without exposure to ANPD fines on the synthetic output, and without having to process personal data for every new question.

ANPD had not published specific guidance on synthetic data by May 2026, but the position of the European Data Protection Board aligns with the interpretation that data which do not allow identification of a natural person fall outside the scope of personal-data regulation.


Why synthetic data are privacy-first by design


What makes synthetic data privacy-first?

Four structural properties: they do not contain references to real individuals; they cannot be reversed to identify people; they do not require individual consent for research use; and they eliminate the risk of leaking sensitive personal data in the output.

  1. No reference to real individuals: the generated personas do not correspond to any person. There is no path from a synthetic persona back to a real individual.

  2. Irreversibility: the synthetic generation process is not reversible. There is no original record to reconstruct from the output.

  3. No individual consent for output: because the output is not personal data, the research workflow does not require consent from synthetic respondents.

  4. No personal-data leakage in the output: teams can share, store, and consult the synthetic output without exposing sensitive personal information.


Market research with real data vs. synthetic data: what changes for marketing teams


Criterion

Research with real data

Research with synthetic data

Consent required

Yes, mandatory legal basis under Art. 7 LGPD

No, output outside LGPD scope

DPO approval

Required before each project

Not required for the synthetic output

Risk of data leakage

Exists whenever personal data are processed

No leakage of personal data in the output

Reuse across teams

Restricted by purpose and consent

Broader reuse of synthetic insights

Time to start

Depends on collection and approvals

Immediate after persona setup

Legal exposure

Higher operational and regulatory risk

Lower exposure in the synthetic output

Scalability

Limited by recruitment and privacy workflow

High scalability after persona creation

Sharing results

Depends on purpose limitation

Can circulate as synthetic insight


The difference between anonymized data and synthetic data

What is the difference between anonymized data and synthetic data for LGPD purposes?

Anonymized data start from real personal data and go through de-identification, with residual risk of re-identification. Synthetic data are artificially generated without relying on individual real records. LGPD does not apply to synthetic data when no natural person can be identified.


Criterion

Anonymized data

Synthetic data

Origin

Real personal data that were de-identified

Artificially generated by algorithms

Re-identification risk

Residual risk

Zero in the output because no real record is present

LGPD applicability

May return if re-identification is possible

Outside scope when not linked to an identifiable person

Governance effort

Requires controls against re-identification

Focuses on input governance and model process


LGPD vs GDPR: what changes for AI market research

Is Galaxies’ synthetic-data approach compatible with European GDPR?

Yes. Both Brazil’s LGPD and Europe’s GDPR exclude from their definitions of personal data information that does not refer to an identified or identifiable natural person. For companies with international operations, the same approach supports research workflows with lower privacy exposure.


Aspect

LGPD (Brazil)

GDPR (Europe)

Definition of personal data

Art. 5, I: identified or identifiable person

Art. 4(1): identified or identifiable person

Synthetic data

Outside scope when not personal data

Outside scope when not personal data

Consent requirement for synthetic output

No

No

Input data obligations

Still apply to real input data

Still apply to real input data

Governance focus

Legal basis and DPA for input

Lawful basis and processing controls for input


Frequently asked questions

Are synthetic data legal in Brazil?

Yes. Synthetic data are not personal data under Art. 5, I of the LGPD because they do not refer to an identified or identifiable natural person. There is no collection or processing of personal data in the synthetic output generated by Galaxies, eliminating individual consent requirements for that output.

Does AI market research require respondent consent?

With synthetic data, no. Because generated personas do not correspond to real individuals, there is no personal-data processing that requires consent. Input data follow the appropriate legal bases, but the synthetic output is outside the LGPD scope.

What is the difference between anonymized data and synthetic data for LGPD?

Anonymized data start from real personal data with residual re-identification risk. Synthetic data are artificially generated without individual records. LGPD does not apply to synthetic data when the output cannot identify a person, offering stronger regulatory protection than anonymization.

How does Galaxies handle input data provided by the client?

Input data are processed under an appropriate legal basis documented in contract with a DPA. The platform does not store third-party personal data beyond what is necessary to generate the model and applies anonymization before AI processing.

How should the synthetic-data approach be presented to the company DPO?

The central argument is legal and direct: the Galaxies platform output is not personal data under Art. 5, I of the LGPD. Input data remain subject to LGPD and are processed under a documented legal basis in a DPA. DPO approval focuses on the input process, not each use of the generated personas.

With the end of third-party cookies, how can consumer research be done at scale?

Synthetic data are one structural answer to the post-cookie scenario. Because they do not depend on individual tracking or navigation data, the methodology is not affected by third-party cookie depreciation in Chrome or by changes in platform privacy policies.

Does AI market research require legal approval before every project?

With synthetic data, no. Legal approval applies to the input process and is resolved once in the contract, not project by project. The synthetic output can be used, shared across departments, and stored without additional legal or DPO approval for each use.


Galaxies