Daniel Victorino


Data does not speak. But it can tell stories when we listen with method.
Recently, I have been working actively on the production of Synthetic Personas at Galaxies. I made many mistakes, learned a lot, and refined many techniques and methods in data science. In academia, we are trained to interpret the world with rigor. In the corporate world, the challenge is different: transforming rigor into action.
That is why I write this article: to share learnings and provoke reflection. This brief text asks: how can we use statistics, machine learning, and generative models to transform behavioral abstractions into strategic assets — Synthetic Personas?
At Galaxies, we constantly face the challenge of deeply understanding diverse audiences, whether users of digital platforms, research respondents, or consumers of innovative products.
Without stereotypes or guesswork
There is a certain quiet beauty and elegance in the mathematics behind audience segmentation. At Galaxies, our approach to building personas does not start from stereotypes or guesswork. It comes from the convergence of empirical data, sociological and economic perspectives, and statistical modeling.
At Galaxies, clustering is more than a technique; it is applied epistemology. For that reason, we do not cling to a single algorithm. As statistician George E. P. Box famously said, “All models are wrong, but some are useful.” We therefore use different unsupervised algorithms to understand each case.
Each model presents different trade-offs, and part of human creativity is understanding and selecting the best one for each case. Clustering algorithms do not “see” users; they organize vectors, numbers, in a space of variables that were carefully selected, transformed, and interpreted.
This work is not automated. It is curated
The process is anchored in an adaptation of the CRISP-DM method, which acts as a structuring framework for critical analysis. It helps ensure clusters are interpretable, reproducible, and above all actionable. Interpreting these groups is not trivial: it requires both technical rigor and contextual sensitivity.
This is where generative models enter. By combining multivariate statistics and different LLMs, we translate statistical results into synthetic narratives. They are not only profiles; they are sociotechnical constructions that integrate observed data with textual inference.
This approach allows us to:
Make the complexity of large datasets legible.
Translate patterns into strategic knowledge.
Create mediation between the world of data and the world of action.
In this context, clustering is almost a philosophical device: it structures how we think, infer, and decide. To understand where both techniques meet, the LLM works like the five senses of the persona, translating what it sees, hears, consumes, and expresses. But clustering is what gives it structure.
Perhaps my greatest challenge as a data scientist applied to business, in the context of Synthetic Personas, is epistemological: how do we transform measures into meanings? That is exactly what we seek to do at Galaxies. More than data science, this is computational social science.
To conclude, I can state with a high degree of confidence that data, when treated with technical responsibility and sociological sensitivity, does not merely inform: it illuminates. And this bridge between statistical rigor and critical interpretation is what we seek to build.
Galaxies


