Publications
* denotes equal contribution.
2026
- Textual Supervision Enhances Geospatial Representations in Vision-Language ModelsMarcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, and 5 more authorsIn International Conference on Machine Learning, 2026
Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation and spatial reasoning. In this work, we analyze the geospatial representations acquired by three model families: vision-only architectures (e.g., ViT), vision-language models (e.g., CLIP), and large-scale multimodal foundation models (e.g., LLaVA, Qwen, and Gemma). By evaluating across image clusters, including people, landmarks, and everyday objects, grouped based on the degree of localizability, we reveal systematic gaps in spatial accuracy and show that textual supervision enhances the learning of geospatial representations. Our findings suggest the role of language as an effective complementary modality for encoding spatial context and multimodal learning as a key direction for advancing geospatial AI.
- Mapping Emerging Climate Misinformation Playbooks in the Global SouthMarcelo Sartori Locatelli*, Wenchao Dong*, Pedro Loures Alzamora, and 4 more authorsIn ACM Conference on Fairness, Accountability and Transparency, 2026Selected for presentation at the ICRN Forum 2026
Climate misinformation continues to erode support for climate action, a challenge that is especially acute in the Global South, where high climate vulnerability intersects with development pressures. In rapidly evolving digital ecosystems, misinformation adapts to platform incentives, shifting from overt rejection of climate science toward more subtle narratives that contest proposed solutions. This study integrates large-scale platform data with qualitative content analysis to examine how information systems shape contemporary climate discourse. Using a dataset of 226,775 climate-related YouTube videos from Brazil (2019-2025), we identify two dominant misinformation strategies: traditional denial that disputes scientific evidence and an emerging “new denial” that accepts climate change while undermining mitigation and adaptation policies. We find a pronounced transition to solution-focused narratives that target renewable energy, climate governance, and environmental advocates. New denial content is produced by a wider array of actors, attracts higher engagement, and employs more sophisticated persuasive techniques. These patterns disproportionately affect regions already facing structural inequities and bring broader concerns about platform accountability in unequal information environments and suggest the need for governance approaches capable of addressing new denial, a rapidly adapting form of harmful content that often evades existing moderation policies.
- Characterizing AI Manipulation Risks in Brazilian YouTube Climate DiscourseWenchao Dong*, Marcelo Sartori Locatelli*, Virgilio Almeida, and 1 more authorIn AAAI Conference on Artificial Intelligence, 2026Selected for presentation at the ICRN Forum 2026
Climate change poses a global threat to public health, food security, and economic stability. Addressing it requires evidence-based policies and a nuanced understanding of how the threat is perceived by the public, particularly within visual social media, where narratives quickly evolve through voices of individuals, politicians, NGOs, and institutions. This study investigates climate-related discourse on YouTube within the Brazilian context, a geopolitically significant nation in global environmental negotiations. Through three case studies, we examine (1) which psychological content traits most effectively drive audience engagement,(2) the extent to which these traits influence content popularity, and (3) whether such insights can inform the design of persuasive synthetic campaigns such as climate denialism using recent generative language models. Another contribution of this work is the release of a large publicly available dataset of 226K Brazilian YouTube videos and 2.7 M user comments on climate change. The dataset includes fine-grained annotations of persuasive strategies, theory of mind categorizations in user responses, and typologies of content creators. This resource can help support future research on digital climate communication and the ethical risk of algorithmically amplified narratives and generative media.
2025
- Characterizing Persuasion Patterns in Climate Discourse on Brazilian Portuguese YouTube VideosWenchao Dong*, Marcelo Sartori Locatelli*, Virgilio Almeida, and 1 more authorIn ACM International Conference on Information Technology for Social Good (GoodIT), 2025
While climate change remains a critical global challenge, research on online climate discourse has been constrained to English-language content, hindering a comprehensive understanding of its global dynamics. This study analyzes 227,159 Brazilian Portuguese YouTube videos to identify persuasive characteristics of climate communication and key discussion themes and players. We find that persuasion strategies rooted in emotional appeal, moral framing, and reciprocity are positively correlated with audience engagement, while logic-based arguments, statistical evidence, and authority narratives exhibit negative influences. We identified 25 key topics and 6 major channel themes, representing the heterogeneous nature of online climate discussion. We discuss the implications of our findings, highlight the potential risks of public opinion manipulation using AI, and outline future research directions.
- Politicization During the 2024 United States Presidential ElectionsMarcelo Sartori Locatelli*, Matheus Prado Miranda*, Wagner Meira Jr., and 1 more authorIn International Conference on Advances in Social Networks Analysis and Mining (ASONAM), 2025
Topic shifts occur naturally during conversations when a person changes the subject to one different from the original. This phenomenon may be extremely meaningful, being useful for modeling dialogues and measuring information distortion during a conversation, among other tasks. In this work, we explore topic shifts as a metric for the measurement of politicization during the 2024 U.S. presidential elections, using YouTube news as a case study. We find evidence of politicization during the studied period as over 69% of non-political videos contain at least one political comment. This politicization increases as we get closer to the date of the election, with videos from right-leaning channels having over 40% of comments being political in the week of the elections. We also identify topics relating to immigration to be the most politicized, as commenters discuss the government’s stance on immigration, aggravated by displays of xenophobia, exemplifying the dangers that come with politicization.
- From Inclusion to Contention: Analyzing DEI and “Woke” Narratives on RedditMarcelo Sartori Locatelli, Arthur Salles Da Costa, Victor Thomé, and 2 more authorsIn International Conference on Advances in Social Networks Analysis and Mining (ASONAM), 2025
Diversity, Equity, and Inclusion (DEI) policies have recently become extremely controversial, with many companies vowing to end their support. This has led to mixed reactions online. This was intensified by the ongoing “woke” vs “anti-woke” culture war. Both groups defend and consume content that aligns with their ideologies. In this context, understanding the discourse surrounding these issues online is essential, as such movements have the potential to lead to real-world harm. For this reason, we conduct a large-scale study around the DEI and “woke” discussion on the Reddit platform from 2020–2024, finding that it has grown significantly during the studied period, spreading across a large variety of seemingly unrelated topics. Finally, we note that the discourse has become increasingly polarized, with a growing trend of toxicity and negative sentiments, coupled with changes in the meaning of the terms “woke” and DEI on the platform. These findings have important implications for public policy related to social issues.
- Analyzing Political Discourse on Discord during the 2024 US Presidential ElectionArthur Buzelin*, Pedro Robles Dutenhefner*, Marcelo Sartori Locatelli*, and 8 more authorsIn Proceedings of the 17th ACM Web Science Conference 2025, 2025
Social media networks have amplified the reach of social and political movements, but most research focuses on mainstream platforms such as X, Reddit, and Facebook, overlooking Discord. As a rapidly growing, community-driven platform with optional decentralized moderation, Discord offers unique opportunities to study political discourse. This study analyzes over 30 million messages from political servers on Discord discussing the 2024 U.S. elections. Servers were classified as Republican-aligned, Democratic-aligned, or unaligned based on their descriptions. We tracked changes in political conversation during key campaign events and identified distinct political valence and implicit biases in semantic association through embedding analysis. We observed that Republican servers emphasized economic policies and Democratic servers focusing on equality-related and progressive causes. Furthermore, we detected an increase in toxic language, such as sexism, in Republican-aligned servers after Kamala Harris’s nomination. These findings provide a first look at political behavior on Discord, highlighting its growing role in shaping and understanding online political engagement.
- AI and Climate Change Discourse: What Opinions Do Large Language Models Present?Marcelo Sartori Locatelli, Pedro Dutenhefner, Arthur Buzelin, and 8 more authorsIn Proceedings of the 2nd Workshop on Natural Language Processing Meets Climate Change (ClimateNLP 2025), 2025
Large Language Models (LLMs) are increasingly used in applications that shape public discourse, yet little is known about whether they reflect distinct opinions on global issues like climate change. This study compares climate change-related responses from multiple LLMs with human opinions collected through the People’s Climate Vote 2024 survey. We compare country and LLM”s answer probability distributions and apply Exploratory Factor Analysis (EFA) to identify latent opinion dimensions. Our findings reveal that while LLM responses do not exhibit significant biases toward specific demographic groups, they encompass a wide range of opinions, sometimes diverging markedly from the majority human perspective.
- Evolutionary Bias Identification with EmbeddingsArthur Buzelin, Yan Aquino, Victoria Estanislau, and 11 more authorsIn International Conference on the Applications of Evolutionary Computation (Part of EvoStar), 2025
This paper introduces EBIE (Evolutionary Bias Identification with Embeddings), a new method to help tackle algorithmic bias in natural language processing (NLP) tasks. The method leverages the powerful representation of word embeddings through an evolutionary algorithm, focusing on classification tasks. EBIE monitors shifts in individual embedding dimensions over generations and, by tracking these dimensional changes, identifies which parts of the embedding are most responsive to changes performed by genetic operations. These insights reveal critical features that influence model decisions and expose latent biases embedded within NLP classifiers. Through correlation analysis between individual tokens and classification scores, EBIE uncovers systematic biases in model behavior, such as reliance on stereotypical markers and neglect of nuanced expressions. By uncovering these tendencies, our methodology provides actionable insights to refine model training, enhance fairness, and improve robustness. Its flexibility ensures broad applicability across various NLP tasks, offering a powerful and versatile framework for developing more equitable and transparent machine learning systems.
- Exploring Brazilian TikTok and YouTube Shorts: A Public Dataset for Video CharacterizationTomas Lacerda, Marcelo Sartori Locatelli, Igor Costa, and 3 more authorsIn Brazilian Workshop on Social Network Analysis and Mining (BraSNAM), 2025
Short video platforms have garnered significant attention in recent years, with discussions ranging from concerns about inappropriate content and addiction to strategies for maximizing user engagement and screen time. Despite the large user base and growing relevance of these platforms, there is still a notable lack of comprehensive datasets focused on broad recommendation and moderation. This is especially true for TikTok, where API access is limited and collecting unbiased data is challenging. In this collection, we present a diverse and rich dataset from YouTube Shorts and TikTok’s main feeds in Brazil, comprising over 35,000 videos. The dataset includes detailed engagement statistics, extensive video metadata, over 170,000 keyframes for visual analysis, and Safe Search API assessments for each keyframe. This rich resource fills a critical data gap, offering valuable tools for research on content categorization, user behavior analysis, and platform engagement strategies.
- Emerging Digital Spaces: The Case of Bluesky’s Rise During 2024 Brazilian ElectionsArthur Buzelin, Pedro Bento, Yan Aquino, and 8 more authorsIn Brazilian Workshop on Social Network Analysis and Mining (BraSNAM), 2025
This paper examines political discourse on Bluesky during Brazil’s 2024 municipal elections, following the unexpected Twitter ban in the country. Using keyword-based data collection and embedding-based sentiment analysis (WEAT), we analyze user perceptions and their evolution across electoral phases. Our findings reveal significant shifts in candidate popularity, closely tied to electoral outcomes and broader ideological trends, with clear patterns of polarization. Notably, candidate sentiment fluctuated between election rounds, reflecting both victories and defeats. This study highlights Bluesky’s emerging role in digital political engagement when mainstream platforms are disrupted.
2024
- Examining the behavior of llm architectures within the framework of standardized national exams in brazilMarcelo Sartori Locatelli, Matheus Prado Miranda, Igor Joaquim Silva Costa, and 8 more authorsIn Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2024
The Exame Nacional do Ensino Médio (ENEM) is a pivotal test for Brazilian students, required for admission to a significant number of universities in Brazil. The test consists of four objective high-school level tests on Math, Humanities, Natural Sciences and Languages, and one writing essay. Students’ answers to the test and to the accompanying socioeconomic status questionnaire are made public every year (albeit anonymized) due to transparency policies from the Brazilian Government. In the context of large language models (LLMs), these data lend themselves nicely to comparing different groups of humans with AI, as we can have access to human and machine answer distributions. We leverage these characteristics of the ENEM dataset and compare GPT-3.5 and 4, and MariTalk, a model trained using Portuguese data, to humans, aiming to ascertain how their answers relate to real societal groups and what that may reveal about the model biases. We divide the human groups by using socioeconomic status (SES), and compare their answer distribution with LLMs for each question and for the essay. We find no significant biases when comparing LLM performance to humans on the multiple-choice Brazilian Portuguese tests, as the distance between model and human answers is mostly determined by the human accuracy. A similar conclusion is found by looking at the generated text as, when analyzing the essays, we observe that human and LLM essays differ in a few key factors, one being the choice of words where model essays were easily separable from human ones. The texts also differ syntactically, with LLM generated essays exhibiting, on average, smaller sentences and less thought units, among other differences. These results suggest that, for Brazilian Portuguese in the ENEM context, LLM outputs represent no group of humans, being significantly different from the answers from Brazilian students across all tests.
- Topic Shifts as a Proxy for Assessing Politicization in Social MediaMarcelo Sartori Locatelli, Pedro Calais, Matheus Prado Miranda, and 4 more authorsIn Proceedings of the International AAAI Conference on Web and Social Media, 2024
Politicization is a social phenomenon studied by political science characterized by the extent to which ideas and facts are given a political tone. A range of topics, such as climate change, religion and vaccines has been subject to increasing politicization in the media and social media platforms. In this work, we propose a computational method for assessing politicization in online conversations based on topic shifts, i.e., the degree to which people switch topics in online conversations. The intuition is that topic shifts from a non-political topic to politics are a direct measure of politicization – making something political, and that the more people switch conversations to politics, the more they perceive politics as playing a vital role in their daily lives. A fundamental challenge that must be addressed when one studies politicization in social media is that, a priori, any topic may be politicized. Hence, any keyword-based method or even machine learning approaches that rely on topic labels to classify topics are expensive to run and potentially ineffective. Instead, we learn from a seed of political keywords and use Positive-Unlabeled (PU) Learning to detect political comments in reaction to non-political news articles posted on Twitter, YouTube, and TikTok during the 2022 Brazilian presidential elections. Our findings indicate that all platforms show evidence of politicization as discussion around topics adjacent to politics such as economy, crime and drugs tend to shift to politics. Even the least politicized topics had the rate in which their topics shift to politics increased in the lead up to the elections and after other political events in Brazil – an evidence of politicization.
2023
- Análise Temporal e Espacial dos Casos de Covid-19 nas Regiões Geográficas Imediatas do BrasilPedro Loures Alzamora, Daniel Victor Ferreira, Isadora Cristina Matos Rodrigues, and 8 more authorsHygeia: Revista Brasileira de Geografia Médica e da Saúde, 2023
O geoprocessamento de dados e as análises espaciais são ferramentas importantes para o estudo de fenômenos como a disseminação de doenças pelo território e ao longo do tempo. O objetivo deste estudo é investigar, utilizando a Análise Exploratória de Dados Espaciais (AEDE), as alterações nos padrões de distribuição geográfica da Covid-19 no Brasil em dois períodos distintos da pandemia: (i) entre abril e agosto de 2020; e (ii) entre novembro de 2020 e março de 2021. Para tanto, as estatísticas I de Moran e LISA foram aplicadas aos dados referentes a três indicadores epidemiológicos: casos acumulados, novos casos e letalidade da doença. Os resultados encontrados e as visualizações propostas apresentam uma perspectiva ampla sobre a variação nos casos de Covid-19 nas regiões brasileiras e colaboram para um melhor entendimento sobre as dinâmicas epidemiológicas no Brasil no primeiro ano da pandemia de Covid-19.
2022
- Characterizing vaccination movements on YouTube in the United States and BrazilMarcelo Sartori Locatelli*, Josemar Caetano*, Wagner Meira Jr, and 1 more authorIn Proceedings of the 33rd ACM Conference on Hypertext and Social Media, 2022
In the context of COVID-19 pandemic, social networks such as Facebook, Twitter, YouTube and Instagram stand out as important sources of information. Among those, YouTube, as the largest and most engaging online media consumption platform, has a large influence in the spread of information and misinformation, which makes it important to study how the platform deals with the problems that arise from disinformation, as well as how its users interact with different types of content. Considering that United States (USA) and Brazil (BR) are two countries with the highest COVID-19 death tolls, we asked the following question: What are the nuances of vaccination campaigns in the two countries? With that in mind, we engage in a comparative analysis of pro and anti-vaccine movements on YouTube. We also investigate the role of YouTube in countering online vaccine misinformation in USA and BR. For this means, we monitored the removal of vaccine related content on the platform and also applied various techniques to analyze the differences in discourse and engagement in pro and anti-vaccine ”comment sections”. We found that American anti-vaccine content tend to lead to considerably more toxic and negative discussion than their pro-vaccine counterparts while also leading to 18% higher user-user engagement, while Brazilian anti-vaccine content was significantly less engaging. We also found that pro-vaccine and anti-vaccine discourses are considerably different as the former is associated with conspiracy theories (e.g. ccp), misinformation and alternative medicine (e.g. hydroxychloroquine), while the latter is associated with protective measures. Finally, it was observed that YouTube content removals are still insufficient, with only approximately 16% of the anti-vaccine content being removed by the end of the studied period, with the United States registering the highest percentage of removed anti-vaccine content(34%) and Brazil registering the lowest(9.8%).
- A COVID-19 no Twitter: correlacionando vocabulário com agravamento e atenuação da pandemia no BrasilPedro Loures Alzamora, Marcelo Sartori Locatelli, Marcelo Ganem, and 8 more authorsIn Brazilian Workshop on Social Network Analysis and Mining (BraSNAM), 2022
O presente estudo busca caracterizar o primeiro ano da pandemia de COVID-19 no Brasil como um fenômeno social por meio da análise da correlação entre o agravamento/atenuação da pandemia e o vocabulário utilizado no Twitter nas semanas que precedem essas variações. Entre outros resultados, observou-se que termos politicamente motivados e com teor negativo são mais prevalentes nas semanas que precedem o aumento do número de casos/mortes, ao passo que o uso de termos relacionados a conteúdos midiáticos (internet, música, televisão) é intensificado nas semanas que antecedem a queda da quantidade de casos/mortes. Tais resultados sugerem a possibilidade de utilização do método aqui introduzido para a análise de fenômenos sociais a partir de dados computacionalmente leves e totalmente anonimizados provenientes de redes sociais online.
- Correlations between web searches and COVID-19 epidemiological indicators in BrazilMarcelo Sartori Locatelli, Evandro LT Cunha, Janaı́na Guiginski, and 8 more authorsBrazilian Archives of Biology and Technology, 2022
COVID-19 rapidly spread across the world in an unprecedented outbreak with a massive number of infected and fatalities. The pandemic was heavily discussed and searched on the internet, which generated big amounts of data related to it. This led to the possibility of attempting to forecast coronavirus indicators using the internet data. For this study, Google Trends statistics for 124 selected search terms related to pandemic were used in an attempt to find which keywords had the best Spearman correlations with a lag, as well as a forecasting model. It was found that keywords related to coronavirus testing among some others, such as “I have contracted covid”, had high correlations (≥0.7) with few weeks of lag (≤4 weeks). Besides that, the ARIMAX model using those keywords had promising results in predicting the increase or decrease of epidemiological indicators, although it was not able to predict their exact values. Thus, we found that Google Trends data may be useful for predicting outbreaks of coronavirus a few weeks before they happen, and may be used as an auxiliary tool in monitoring and forecasting the disease in Brazil.