Enhancing Image Generation with GANs: The Role of Mutual Information in Optimizing Generative Models

Sitaresmi Wahyu Handani, Pintusorn Suttiponpisarn, Gwan-yen Lin, Ruey-Feng Chang

Abstract


Generative Adversarial Networks (GANs) have become a prominent approach for image generation; however, they often suffer from training instability, mode collapse, and limited controllability of generated outputs. This study investigates the role of mutual information in improving generative modeling through a comparative analysis of Vanilla GAN, Conditional GAN (CGAN), and InfoGAN. Experiments were conducted using two image datasets with different levels of complexity, namely MNIST and Anime Face, under comparable training configurations. The evaluation focused on training behavior, convergence characteristics, generated image quality, and latent representation learning. The results revealed notable differences among the evaluated models. Vanilla GAN exhibited unstable convergence behavior at higher training epochs, while CGAN provided conditional control over generated outputs but did not fully mitigate training instability. In contrast, InfoGAN maintained more balanced generator and discriminator loss dynamics and produced visually consistent outputs across both datasets. Furthermore, latent code manipulation experiments showed that InfoGAN learned more structured and interpretable latent representations, enabling controllable feature variation in generated images. These findings indicate that incorporating mutual information improves representation learning, controllability, and training stability in GAN-based image generation. This study provides an empirical comparison of GAN, CGAN, and InfoGAN under a unified experimental framework and demonstrates that mutual information regularization contributes to improved training stability, controllable generation, and more interpretable latent representations. The findings highlight the potential of information-theoretic regularization for enhancing generative modeling performance.

Keywords


Generative Adversarial Networks (GANs); Mutual Information; InfoGAN; Image Generation; Latent Representation

Full Text:

Link Download

References


Anime Face Dataset. (2019).

Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., & Abbeel, P. (2016). InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets. https://arxiv.org/abs/1606.03657.

Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative Adversarial Networks.

Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2020). Generative adversarial networks. Communications of the ACM, 63(11), 139–144. https://doi.org/10.1145/3422622

Handani, S. W., & Chang, R. F. (2025). A Comprehensive Review of Dual and Multi-Generator Architectures in GAN-Based Image-to-Image Translation. Proceedings - 2025 9th International Conference on Information Technology, Information Systems and Electrical Engineering, ICITISEE 2025, 439–444. https://doi.org/10.1109/ICITISEE68184.2025.11355114

Kim, J.-H., Kim Jiyoung Lee, Y., & Min Yoo, K. (2022). Mutual Information Divergence: A Unified Metric for Multimodal Generative Models. https://github.com/naver-ai/mid.metric.

Lecun, Y., Bottou, E., Bengio, Y., & Haffner, P. (1998). Gradient-Based Learning Applied to Document Recognition.

Li, W., Liang, Z., Neuman, J., Chen, J., & Cui, X. (2021). Multi-generator GAN learning disconnected manifolds with mutual information. Knowledge-Based Systems, 212. https://doi.org/10.1016/j.knosys.2020.106513

Mirza, M., & Osindero, S. (2014). Conditional Generative Adversarial Nets. http://arxiv.org/abs/1411.1784

Russakoff, D. B., Tomasi, C., Rohlfing, T., & Maurer, C. R. (2004). LNCS 3023 - Image Similarity Using Mutual Information of Regions.

Singh, N. K., & Raza, K. (2021). Medical Image Generation Using Generative Adversarial Networks: A Review. In Studies in Computational Intelligence (Vol. 932, pp. 77–96). Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/978-981-15-9735-0_5

Zhai, Y. K., Long, Z. H., Pan, W. F., & Chen, C. L. P. (2024). Mutual Information Compensation for High-Fidelity Image Generation with Limited Data. IEEE Signal Processing Letters, 31, 2145–2149. https://doi.org/10.1109/LSP.2024.3439131

Zhang, Z., Li, M., & Yu, J. (2018). On the convergence and mode collapse of GAN. SIGGRAPH Asia 2018 Technical Briefs, 1–4. https://doi.org/10.1145/3283254.3283282




DOI: http://dx.doi.org/10.35671/telematika.v19i2.3304

Refbacks

  • There are currently no refbacks.


 



Indexed by:

   

Telematika
ISSN: 2442-4528 (online) | ISSN: 1979-925X (print)
Published by : Universitas Amikom Purwokerto
Jl. Let. Jend. POL SUMARTO Watumas, Purwonegoro - Purwokerto, Indonesia


Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License .