Daniel Victorino


Galaxies Synthetic Personas are created in five stages: input-data validation, Machine Learning clustering, LLM generation, post-model statistical validation, and activation in the platform. The process uses only real data provided by the client, collected by Galaxies, or sourced from authorized partners.
The question every client asks first
When someone first hears that it is possible to generate hundreds of consumer profiles with Artificial Intelligence, the natural reaction is skepticism. “Isn’t this just ChatGPT inventing answers?” “How can this persona represent my real customer if it is synthetic?”
These are good questions. And they deserve specific answers, not generic ones.
The difference between a quality Synthetic Persona and AI generating random answers is exactly the creation methodology. Galaxies developed a five-stage process with quality controls at every phase, and this article opens that process step by step.
The fundamental principle: real data as the starting point
Galaxies Synthetic Personas never start from zero. They always start from real research data.
Why does this matter? Because it is the difference between a persona that represents your real consumer and a persona that represents the average consumer imagined by a language model trained on internet data.
The platform accepts three types of input data:
Client data: previous quantitative research (spreadsheets, CSV), qualitative interview transcripts, CRM data, and NPS results.
Data collected by Galaxies: questionnaires applied by the platform for the specific project.
Authorized partner data: datasets with documented origin, consent, and appropriate governance.
What is not accepted or recommended: internet-scraped data, purchased list data, and data without clear provenance. This restriction is a technical and ethical choice, and it directly affects the quality and LGPD compliance of the result.
The five stages of creating a synthetic persona
Stage 1Pre-Model Statistical Validation |
Before any processing, the platform performs a complete audit of the input data. The goal is to ensure that the dataset has enough quality to generate reliable personas.
This stage checks variable distribution, response consistency, outlier presence, minimum respondent volume per segment, and the integrity of categorical and numerical data.
If the dataset has problems — insufficient sample, evident selection bias, internal inconsistencies — the system flags them before continuing. This avoids what data scientists call “garbage in, garbage out”: if the input data is poor, the generated output will be poor too.
Only after approval in pre-model validation do the data move to the next stage.
Stage 2Machine Learning Clustering |
With validated data, Machine Learning algorithms group real respondents into homogeneous clusters. Each cluster represents a distinct consumer segment with similar characteristics.
Clustering variables include demographic data, attitudinal data, behavioral data, and lifestyle data: age, income, region, education, values, beliefs, preferences, purchase habits, usage frequency, preferred channels, and life context.
The critical point here is representativeness. The model is calibrated to ensure that minority groups are proportionally represented, not only the most frequent profiles. Research that only represents the average consumer is not enough for strategic decisions.
The result of this stage is a set of well-defined groups, with clear boundaries, which will serve as the basis for persona generation in the next stage.
Stage 3Persona Generation with LLM |
This is where the large language model enters. For each cluster generated in the previous stage, the LLM creates personas using real customer data as the behavioral anchor.
This is fundamentally different from asking ChatGPT to invent a consumer profile. The LLM is not imagining; it is synthesizing. Persona answers are anchored in the real statistical patterns of the group they represent.
The result is a persona with coherent identity: it talks about preferences, objections, habits, and needs in a way that is consistent with what the real group demonstrated in research. When a platform user asks the persona a question, the response follows that behavioral anchor.
After this stage, the persona can already answer questions in natural language. But it is not yet released for use; the validation stage is still missing.
Stage 4Post-Model Statistical Validation |
This is the stage that differentiates the Galaxies methodology from any generic AI persona solution. Before making personas available to the client, the platform submits results to a battery of tests using the DeepEval framework.
This framework is customized for Galaxies and evaluates results across five dimensions:
Relevance: do persona answers actually address what was asked? (76.5% approval, the most demanding dimension by design).
Bias: is there bias or differential treatment in responses based on demographic characteristics? (91.3% approval).
Hallucination control: does the answer stay grounded in the persona and input data? (97.8% approval).
Persona fidelity: does the answer remain consistent with the persona profile? (95.1% approval).
Toxicity: does the model avoid toxic or unsafe responses? (100% approval).
Only personas that pass these tests within acceptable parameters are released. Personas outside the expected result are submitted to calibration and refinements until they reach the minimum acceptable levels before being released.
97.8% Hallucination control (DeepEval) | 95.1% Persona fidelity (DeepEval) | 100% Absence of toxicity (DeepEval) | 91.3% Bias control (DeepEval) |
How does the accuracy-analysis process work?
Galaxies asks new questions — different from the original training questions, but within the same segment — to organic respondents and Synthetic Personas, then compares how often the answers are identical or point in the same direction.
This process makes the accuracy index of Synthetic Personas generated by the tool exceed 85%.
Stage 5Active Persona in Nexus: Galaxies Lab |
With validation completed, personas are available in Nexus: Galaxies Lab. From this moment, the user can:
Talk to personas: in natural language, as they would in an interview.
Ask open questions: and receive contextualized answers aligned with each persona profile.
Simulate scenarios: price changes, product reformulation, new campaigns, or strategic shifts.
Compare clusters: understand how different segments react to the same stimulus.
Analyze answers by cluster and identify patterns that would not be visible in a single average consumer profile.
Export insights for marketing, product, innovation, and strategy teams.
What makes the Galaxies methodology different from alternatives?
There are other tools in the market that propose generating personas with AI. The difference is in four specific points:
1. Proprietary data, not internet data
Most alternative solutions use internet data, social networks, forums, and reviews to train or feed personas. This creates generic profiles based on public behaviors that may have nothing to do with the client’s actual consumer. Galaxies starts from real project data.
2. Double statistical validation, pre and post
Most solutions do not validate input-data quality and do not test outputs afterward. Galaxies does both, which drastically reduces the risk of delivering personas with bias, inconsistencies, or hallucinations.
3. Methodological transparency
Galaxies’ methodology is explicit and documented. The client has access to details on how data were treated, how clusters were generated, and what validation-test results showed. This is uncommon in the market and essential for teams that need governance.
4. Partnership with Google Cloud and Nvidia
The platform’s technology infrastructure is developed with support from the Google Cloud acceleration program and the Nvidia Inception Program. This ensures scalability, security, and access to advanced processing capabilities available in the market.
How long does it take to create Synthetic Personas?
Total time depends mainly on the input-data collection and validation stage. If the client already has previous research data available, the process is faster. If the dataset needs to be collected from scratch, fieldwork adds time.
Scenario Estimated time | Client has research data available and a low-complexity database (direct upload) 48 to 72 hours | Fast collection through the platform questionnaire (100–300 respondents) 5 to 10 days | Complete study with collection, personas, and validation 2 to 3 weeks | High-complexity project with multiple segments and validation rounds 3 to 4 weeks |
Frequently asked questions
What data are used to create synthetic personas?
Quantitative and qualitative research data provided by the client, collected by Galaxies, or sourced from authorized partners. The platform never uses internet data, ensuring personalization and LGPD compliance.
How long does it take to create synthetic personas?
If the client already has previous research data, personas are ready in 48 to 72 hours. With data collection from scratch, the process takes from 5 days to 3 weeks — still significantly faster than any equivalent traditional methodology.
How does Galaxies guarantee the accuracy of synthetic personas?
Through double statistical validation: Pre-Model Validation for input-data integrity and Post-Model Validation with model validation and control-group accuracy analysis. Average accuracy is 91% compared with real respondents.
Quality is not an accident
Trust in Synthetic Personas does not come from a marketing promise. It comes from a process with quality controls at each stage, audited input data, and tested results before reaching the client.
Understanding this process is what turns legitimate skepticism into justified confidence. And justified confidence is what allows a company to make million-real decisions based on AI-generated data.
The methodology exists. The results are documented. The cases are verifiable. The next step is to see all of this working for your specific use case.
Schedule a technical session with the Galaxies team to see the process live
Galaxies


