Low-Precision Neural Network Acceleration for Sustainable Generative AI on Heterogeneous Edge Hardware

Authors

  • Massimo Jimenez Department of Computer Science, University of Houston, Houston, TX, USA. Author
  • Merk A. Peters Department of Computer Science, Colorado State University, Fort Collins, CO, USA. Author
  • Enzo M. Lane Department of Computer Science, George Mason University, Fairfax, VA, USA. Author

Keywords:

low-precision neural networks; generative artificial intelligence; edge computing; heterogeneous hardware; sustainable AI; quantized inference; system governance

Abstract

Generative artificial intelligence has begun to migrate from centralized cloud facilities toward heterogeneous edge infrastructures, where constrained energy budgets, variable thermal envelopes, and diverse accelerator architectures create urgent pressures for more efficient inference. Low-precision neural network acceleration has emerged as a promising system-level response because it can reduce memory traffic, lower arithmetic energy, and enable larger generative models to operate within tight edge resource boundaries. This paper examines the structural trade-offs involved in deploying quantized generative workloads across heterogeneous edge hardware, placing emphasis on runtime orchestration, architectural specialization, sustainability governance, robustness, fairness, and policy implications. It argues that low-precision acceleration should not be treated solely as a numerical compression technique, but rather as a systemic design principle that reshapes the relationship between model serving infrastructure, energy consumption, and socio-technical accountability. The discussion integrates perspectives from computer architecture, distributed systems, machine learning operations, environmental assessment, and regulatory governance. Through a critical synthesis of recent research and industrial developments, the paper identifies enduring tensions between aggressive quantization and model reliability, between hardware heterogeneity and predictable service quality, and between carbon efficiency and equitable access. It concludes that sustainable generative AI on edge systems requires coordinated advances in hardware-aware compilation, adaptive runtime management, transparent energy accounting, and policy frameworks that reward measurable efficiency without exacerbating digital divides.

References

1. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

2. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645–3650).

3. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.

4. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2704–2713).

5. Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). LLM.int8(): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems, 35, 30312–30324.

6. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations.

7. Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., & Han, S. (2024). AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. Proceedings of Machine Learning and Systems, 6, 87–100.

8. Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In The International Conference on Learning Representations.

9. Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., ... Jia, Z. (2024). Towards efficient generative large language model serving: A survey from algorithms to systems. arXiv preprint arXiv:2312.15234.

10. Chen, Ce, et al. "JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators." arXiv preprint arXiv:2606.28421 (2026).

11. Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., Shen, H., Wang, L., Hu, Y., Ceze, L., Guestrin, C., & Krishnamurthy, A. (2018). TVM: An automated end-to-end optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (pp. 578–594).

12. Fowers, J., Ovtcharov, K., Papamichael, M., Massengill, T., Liu, M., Lo, D., Alkalay, S., Haselman, M., Adams, L., Ghandi, M., Heil, S., Patel, P., Sapek, A., Weisz, G., Woods, L., Lanka, S., Reinhardt, S. K., Caulfield, A. M., Chung, E. S., & Burger, D. (2018). A configurable cloud-scale DNN processor for real-time AI. In Proceedings of the 45th Annual International Symposium on Computer Architecture (pp. 1–14).

13. Banbury, C., Reddi, V. J., Torelli, P., Holleman, J., Jeffries, N., Kiraly, C., Montino, P., Kanter, D., Ahmed, S., Pau, D., Thakker, U., Torrini, A., Warden, P., Cordaro, J., Di Guglielmo, G., Duarte, J., Gibellini, S., Parekh, V., Tran, H., ... Xuesong, X. (2021). Benchmarking TinyML systems: Challenges and direction. arXiv preprint arXiv:2003.04821.

14. Wang, Y., Wei, G., & Brooks, D. (2019). Benchmarking TPU, GPU, and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701.

15. Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., Chang, G., Aga, F., Huang, J., Bai, C., Gschwind, M., Gupta, A., Ott, M., Melnikov, A., Candido, S., Brooks, D., Chauhan, G., Lee, B., Lee, H., ... Hazelwood, K. (2022). Sustainable AI: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4, 795–813.

16. Luccioni, A. S., Jernite, Y., & Strubell, E. (2024). Power hungry processing: Watts driving the cost of AI deployment? In The 2024 ACM Conference on Fairness, Accountability, and Transparency (pp. 85–99).

17. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).

18. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.

19. Satyanarayanan, M. (2017). The emergence of edge computing. Computer, 50(1), 30–39.

20. European Parliament and Council of the European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence. Official Journal of the European Union.

Downloads

Published

2026-08-07

How to Cite

Low-Precision Neural Network Acceleration for Sustainable Generative AI on Heterogeneous Edge Hardware. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(2). https://www.ijaies.org/index.php/home/article/view/110