Social Identity Reconstruction through Large Language Models: A Computational Approach to Self-Perception Analysis in Online Communities

Authors

  • Yash Dubey School of Computing, Clemson University, Clemson, SC, USA. Author

Keywords:

large language models, social identity, self-perception, online communities, system architecture, fairness, governance, computational social science

Abstract

The proliferation of online communities has transformed how individuals construct, express, and revise social identity. Large language models (LLMs) now offer unprecedented capabilities for analyzing and modeling self-perception narratives at scale, yet their integration into identity-sensitive computational systems raises profound architectural, governance, and ethical challenges. This paper presents a system-level examination of a computational pipeline that ingests user-generated discourse from digital platforms and reconstructs latent identity signals through LLM-based semantic analysis. The discussion emphasizes structural trade-offs in data collection and model deployment, highlighting tensions between accuracy, fairness, interpretability, and privacy. A multi-tier architecture is proposed, coupling federated data governance with modular prompt engineering and human-in-the-loop validation. Robustness is addressed through adversarial stress testing and fairness audits that account for sociolinguistic diversity across communities. Policy implications are examined, including the reconfiguration of identity agency, potential for algorithmic reification of stereotypes, and the need for participatory infrastructure design. By framing identity reconstruction as a socio-technical system, the paper provides a forward-looking perspective on how LLM-driven self-perception analysis can be deployed sustainably and equitably, without undermining the fluid, co-constructed nature of online identity. This work contributes a comprehensive architectural and policy-aware blueprint for researchers and platform designers seeking to harness LLMs for identity research while upholding community values and individual autonomy.

References

1. Bamman, D., Eisenstein, J., & Schnoebelen, T. (2014). Gender identity and lexical variation in social media. Journal of Sociolinguistics, 18(2), 135–160.

2. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics.

3. Hovy, D., & Yang, D. (2021). The importance of modeling social factors of language. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 588–602). Association for Computational Linguistics.

4. Lee, M., Liang, P., & Yang, Q. (2022). Coauthor: Designing a human-AI collaborative writing dataset for exploring language model capabilities. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (pp. 1–19). ACM.

5. Andalibi, N., Haimson, O. L., De Choudhury, M., & Forte, A. (2016). Understanding social media disclosures of sexual abuse through the lenses of support seeking and anonymity. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (pp. 3906–3918). ACM.

6. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., ... & Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210.

7. Blodgett, S. L., Barocas, S., Daumé III, H., & Wallach, H. (2020). Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5454–5476). Association for Computational Linguistics.

8. Birhane, A., Isaac, W., Prabhakaran, V., Díaz, M., Elish, M. C., Gabriel, I., & Mohamed, S. (2022). Power to the people? Opportunities and challenges for participatory AI. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (pp. 1–8). ACM.

9. Wachter, S., Mittelstadt, B., & Floridi, L. (2017). Why a right to explanation of automated decision-making does not exist in the general data protection regulation. International Data Privacy Law, 7(2), 76–99.

10. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). ACM.

11. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59–68). ACM.

12. Lazer, D., Pentland, A., Adamic, L., Aral, S., Barabási, A. L., Brewer, D., ... & Van Alstyne, M. (2009). Computational social science. Science, 323(5915), 721–723.

13. Solanki, D., Hsu, H. M., Zhao, O., Zhang, R., Bi, W., & Kannan, R. (2020, July). The way we think about ourselves. In International Conference on Human-Computer Interaction (pp. 276-285). Cham: Springer International Publishing.

14. Burke, P. J., & Stets, J. E. (2009). Identity theory. Oxford University Press.

15. Tajfel, H., & Turner, J. C. (2004). The social identity theory of intergroup behavior. In J. T. Jost & J. Sidanius (Eds.), Political psychology: Key readings (pp. 276–293). Psychology Press.

16. Marwick, A. E., & boyd, d. (2014). Networked privacy: How teenagers negotiate context in social media. New Media & Society, 16(7), 1051–1067.

17. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1).

18. Zuboff, S. (2019). The age of surveillance capitalism: The fight for a human future at the new frontier of power. PublicAffairs.

19. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., ... & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). ACM.

20. Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., ... & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (pp. 33–44). ACM.

Downloads

Published

2026-06-15

How to Cite

Social Identity Reconstruction through Large Language Models: A Computational Approach to Self-Perception Analysis in Online Communities. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(2). https://www.ijaies.org/index.php/home/article/view/94