Enhancing Image Generation with GANs: The Role of Mutual Information in Optimizing Generative Models

Sitaresmi Wahyu Handani, Pintusorn Suttiponpisarn, Gwan-yen Lin, Ruey-Feng Chang

Abstract


Generative Adversarial Networks (GANs) have become a prominent approach for image generation; however, they often suffer from training instability, mode collapse, and limited controllability of generated outputs. This study investigates the role of mutual information in improving generative modeling through a comparative analysis of Vanilla GAN, Conditional GAN (CGAN), and InfoGAN. Experiments were conducted using two image datasets with different levels of complexity, namely MNIST and Anime Face, under comparable training configurations. The evaluation focused on training behavior, convergence characteristics, generated image quality, and latent representation learning. The results revealed notable differences among the evaluated models. Vanilla GAN exhibited unstable convergence behavior at higher training epochs, while CGAN provided conditional control over generated outputs but did not fully mitigate training instability. In contrast, InfoGAN maintained more balanced generator and discriminator loss dynamics and produced visually consistent outputs across both datasets. Furthermore, latent code manipulation experiments showed that InfoGAN learned more structured and interpretable latent representations, enabling controllable feature variation in generated images. These findings indicate that incorporating mutual information improves representation learning, controllability, and training stability in GAN-based image generation. This study provides an empirical comparison of GAN, CGAN, and InfoGAN under a unified experimental framework and demonstrates that mutual information regularization contributes to improved training stability, controllable generation, and more interpretable latent representations. The findings highlight the potential of information-theoretic regularization for enhancing generative modeling performance.

Keywords


Generative Adversarial Networks (GANs); Mutual Information; InfoGAN; Image Generation; Latent Representation

Full Text:

Link Download

References


Bermano, A. H., Gal, R., Alaluf, Y., Mokady, R., Nitzan, Y., Tov, O., Patashnik, O., & Cohen-Or, D. (2022). State-of-the-Art in the Architecture, Methods and Applications of StyleGAN. Computer Graphics Forum, 41(2), 591–611. https://doi.org/10.1111/CGF.14503

Cao, H., Tan, C., Gao, Z., Xu, Y., Chen, G., Heng, P. A., & Li, S. Z. (2024). A Survey on Generative Diffusion Models. IEEE Transactions on Knowledge and Data Engineering, 36(7), 2814–2830. https://doi.org/10.1109/TKDE.2024.3361474

Chakraborty, T., Reddy K S, U., Naik, S. M., Panja, M., & Manvitha, B. (2024). Ten years of generative adversarial nets (GANs): a survey of the state-of-the-art. Machine Learning: Science and Technology, 5(1), 011001. https://doi.org/10.1088/2632-2153/AD1F77

Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., & Abbeel, P. (2016). InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets. Advances in Neural Information Processing Systems, 29.

Croitoru, F. A., Hondru, V., Ionescu, R. T., & Shah, M. (2023). Diffusion Models in Vision: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), 10850–10869. https://doi.org/10.1109/TPAMI.2023.3261988

Deng, H., Wu, Q., Huang, H., Yang, X., & Wang, Z. (2023). InvolutionGAN: lightweight GAN with involution for unsupervised image-to-image translation. Neural Computing and Applications 2023 35:22, 35(22), 16593–16605. https://doi.org/10.1007/S00521-023-08530-Z

Efatinasab, E., Brighente, A., Donadel, D., Conti, M., & Rampazzo, M. (2025). Towards robust stability prediction in smart grids: GAN-based approach under data constraints and adversarial challenges. Internet of Things, 33, 101662. https://doi.org/10.1016/J.IOT.2025.101662

Golfe, A., del Amor, R., Colomer, A., Sales, M. A., Terradez, L., & Naranjo, V. (2023). ProGleason-GAN: Conditional progressive growing GAN for prostatic cancer Gleason grade patch synthesis. Computer Methods and Programs in Biomedicine, 240, 107695. https://doi.org/10.1016/J.CMPB.2023.107695

Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative Adversarial Networks.

Handani, S. W., & Chang, R. F. (2025). A Comprehensive Review of Dual and Multi-Generator Architectures in GAN-Based Image-to-Image Translation. Proceedings - 2025 9th International Conference on Information Technology, Information Systems and Electrical Engineering, ICITISEE 2025, 439–444. https://doi.org/10.1109/ICITISEE68184.2025.11355114

Isola, P., Zhu, J. Y., Zhou, T., & Efros, A. A. (2017). Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1125-1134).

Kuntalp, M., & Düzyel, O. (2024). A new method for GAN-based data augmentation for classes with distinct clusters. Expert Systems with Applications, 235, 121199. https://doi.org/10.1016/J.ESWA.2023.121199

Lecun, Y., Bottou, E., Bengio, Y., & Haffner, P. (1998). Gradient-Based Learning Applied to Document Recognition.

Li, B., Zhu, Y., Wang, Y., Lin, C. W., Ghanem, B., & Shen, L. (2022). AniGAN: Style-Guided Generative Adversarial Networks for Unsupervised Anime Face Generation. IEEE Transactions on Multimedia, 24, 4077–4091. https://doi.org/10.1109/TMM.2021.3113786

Li, W., Liang, Z., Neuman, J., Chen, J., & Cui, X. (2021). Multi-generator GAN learning disconnected manifolds with mutual information. Knowledge-Based Systems, 212. https://doi.org/10.1016/j.knosys.2020.106513

Li, Y., Chen, J., Zhang, Z., Xie, X., Xu, T., Ma, K., & Zheng, Y. (2022). Beyond mutual information: Generative adversarial network for domain adaptation using information bottleneck constraint. IEEE Transactions on Medical Imaging, 41(3), 595–607. https://doi.org/10.1109/TMI.2021.3117996

Lin, C., Xiong, S., & Chen, Y. (2022). Mutual information maximizing GAN inversion for real face with identity preservation. Journal of Visual Communication and Image Representation, 87, 103566. https://doi.org/10.1016/J.JVCIR.2022.103566

Luo, Y., & Yang, Z. (2024). DynGAN: Solving Mode Collapse in GANs With Dynamic Clustering. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5493–5503. https://doi.org/10.1109/TPAMI.2024.3367532

Lv, F., Liu, G., Wang, Q., Lu, X., Lei, S., Wang, S., & Ma, K. (2022). Pattern Recognition of Partial Discharge in Power Transformer Based on InfoGAN and CNN. Journal of Electrical Engineering & Technology 2022 18:2, 18(2), 829–841. https://doi.org/10.1007/S42835-022-01260-7

Mirza, M., & Osindero, S. (2014). Conditional Generative Adversarial Nets. http://arxiv.org/abs/1411.1784

Najari, S., Salehi, M., Farahbakhsh, R., & Tyson, G. (2023). MidGAN: Mutual information in GAN-based dialogue models. Applied Soft Computing, 148, 110909. https://doi.org/10.1016/J.ASOC.2023.110909

Nayak, A. A., Venugopala, P. S., & Ashwini, B. (2024). A Systematic Review on Generative Adversarial Network (GAN): Challenges and Future Directions. Archives of Computational Methods in Engineering 2024 31:8, 31(8), 4739–4772. https://doi.org/10.1007/S11831-024-10119-1

Park, S., & Shin, Y. G. (2025). A Novel Generator with Auxiliary Branch for Improving GAN Performance. IEEE Transactions on Neural Networks and Learning Systems, 36(3), 5818–5825. https://doi.org/10.1109/TNNLS.2024.3361087

Rao, X., Min, W., Deng, Z., & Liu, M. (2025). Facial expression transformation for anime-style image based on decoder control and attention mask. Signal Processing: Image Communication, 138, 117343. https://doi.org/10.1016/J.IMAGE.2025.117343

Su, Q., Hamed, H. N. A., Isa, M. A., Hao, X., & Dai, X. (2024). A GAN-Based Data Augmentation Method for Imbalanced Multi-Class Skin Lesion Classification. IEEE Access, 12, 16498–16513. https://doi.org/10.1109/ACCESS.2024.3360215

Trinh, L. T., & Hamagami, T. (2024). Latent Denoising Diffusion GAN: Faster Sampling, Higher Image Quality. IEEE Access, 12, 78161–78172. https://doi.org/10.1109/ACCESS.2024.3406535

Yang, L. C., & Lerch, A. (2018). On the evaluation of generative models in music. Neural Computing and Applications 2018 32:9, 32(9), 4773–4784. https://doi.org/10.1007/S00521-018-3849-7

Zhai, Y. K., Long, Z. H., Pan, W. F., & Chen, C. L. P. (2024). Mutual Information Compensation for High-Fidelity Image Generation with Limited Data. IEEE Signal Processing Letters, 31, 2145–2149. https://doi.org/10.1109/LSP.2024.3439131

Zhang, H., Sindagi, V., & Patel, V. M. (2020). Image De-Raining Using a Conditional Generative Adversarial Network. IEEE Transactions on Circuits and Systems for Video Technology, 30(11), 3943–3956. https://doi.org/10.1109/TCSVT.2019.2920407

Zheng, Z., Fan, C., Wang, C., Wang, M., He, X., & He, X. (2026). A GAN integrating CNN and transformer with mutual information and grayscale-based loss functions for modality translation in medical image. Biomedical Signal Processing and Control, 113, 109062. https://doi.org/10.1016/J.BSPC.2025.109062

Zhou, Z., Li, Y., Liu, R., Xu, X., & Yan, Z. (2025). Unsupervised and controllable synthesizing for imbalanced energy dataset based on AC-InfoGAN. Applied Energy, 393, 126107. https://doi.org/10.1016/J.APENERGY.2025.126107




DOI: http://dx.doi.org/10.35671/telematika.v19i2.3304

Refbacks

  • There are currently no refbacks.


 



Indexed by:

   

Telematika
ISSN: 2442-4528 (online) | ISSN: 1979-925X (print)
Published by : Universitas Amikom Purwokerto
Jl. Let. Jend. POL SUMARTO Watumas, Purwonegoro - Purwokerto, Indonesia


Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License .