Cross-Cultural Self-Representation Analysis Using Large Language Models and Social Media Text Mining

Authors

  • Isaac Duncan Department of Computer Science and Engineering, University at Buffalo, Buffalo, NY, USA. Author

Keywords:

large language models, social media mining, cross-cultural analysis, self-representation, system architecture, fairness, governance

Abstract

The proliferation of social media platforms has generated an unprecedented volume of textual data that encodes culturally mediated patterns of self-representation. Understanding how individuals across different cultures construct, perform, and narrate their identities in digital spaces holds profound implications for psychology, computational social science, and the design of globally inclusive artificial intelligence systems. This paper presents a system-level investigation into the deployment of large language models (LLMs) for cross-cultural self-representation analysis at scale. We examine the architectural trade-offs involved in integrating pretrained transformer-based models with social media text mining pipelines, focusing on challenges of linguistic diversity, contextual drift, and representational fairness. The discussion encompasses data infrastructure considerations, including the construction of culturally stratified corpora, preprocessing workflows that balance noise reduction with preservation of vernacular expression, and model adaptation strategies such as parameter-efficient fine-tuning and prompt engineering. We further analyze governance frameworks, ethical risks of cultural flattening, and the sustainability of large-scale inference under real-world deployment constraints. By comparing monolithic versus modular system architectures, we reveal how choices in data filtering, model selection, and evaluation protocols shape the validity of cross-cultural inferences. The paper concludes with policy implications for platform operators, researchers, and regulators, advocating for participatory design processes that embed cultural domain expertise throughout the entire machine learning lifecycle. Throughout, we emphasize that treating culture as a measurable variable rather than an emergent, relational phenomenon can lead to brittle generalizations, and we propose a relational systems approach that foregrounds structural positionality.

References

1. Markus, H. R., & Kitayama, S. (1991). Culture and the self: Implications for cognition, emotion, and motivation. Psychological Review, 98(2), 224–253.

2. Pennebaker, J. W., Mehl, M. R., & Niederhoffer, K. G. (2003). Psychological aspects of natural language use: Our words, our selves. Annual Review of Psychology, 54(1), 547–577.

3. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (pp. 5998–6008).

4. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 4171–4186).

5. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems (pp. 1877–1901).

6. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).

7. Rogers, R. (2020). The social life of data: Controversies in social media data research. In L. Hjorth, H. Horst, A. Galloway, & G. Bell (Eds.), The Routledge Companion to Digital Ethnography (pp. 147–156). Routledge.

8. Boyd, R. L., & Pennebaker, J. W. (2017). Language-based personality: A new approach to personality in a digital world. Current Opinion in Behavioral Sciences, 18, 63–68.

9. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

10. Solanki, D., Hsu, H. M., Zhao, O., Zhang, R., Bi, W., & Kannan, R. (2020, July). The way we think about ourselves. In International Conference on Human-Computer Interaction (pp. 276-285). Cham: Springer International Publishing.

11. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.

12. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59–68).

13. Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., ... & Vayena, E. (2018). AI4People—An ethical framework for a good AI society. Minds and Machines, 28(4), 689–707.

14. Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature Machine Intelligence, 1(9), 389–399.

15. Tufekci, Z. (2014). Big questions for social media big data: Representativeness, validity, and other methodological pitfalls. In Proceedings of the International AAAI Conference on Web and Social Media (pp. 505–514).

16. Olteanu, A., Castillo, C., Diaz, F., & Kıcıman, E. (2019). Social data: Biases, methodological pitfalls, and ethical boundaries. Frontiers in Big Data, 2, 13.

17. Jurafsky, D., & Martin, J. H. (2022). Speech and Language Processing (3rd ed. draft). Prentice Hall.

18. Hofstede, G. (2001). Culture's Consequences: Comparing Values, Behaviors, Institutions, and Organizations Across Nations (2nd ed.). Sage.

19. Ruths, D., & Pfeffer, J. (2014). Social media for large studies of behavior. Science, 346(6213), 1063–1064.

20. Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2), 2053951716679679.

Downloads

Published

2026-06-15

How to Cite

Cross-Cultural Self-Representation Analysis Using Large Language Models and Social Media Text Mining. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(2). https://www.ijaies.org/index.php/home/article/view/93