Skip to main navigation Skip to search Skip to main content

Infusing Multimodal Latent Embedding for Image Deblurring

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Blurred images significantly degrade the performance of vision-based systems. While some blur effects are intentional, such as bokeh, most applications prioritize structural restoration. We propose a novel multimodal framework for image deblurring that infuses high-level semantic features extracted from pretrained multimodal models of CLIP and Stable Diffusion into a two-stage restoration pipeline. Built upon the Adaptive Filter-Based Deblurring Module (AFDM), our approach combines low-level pixel refinement with semantic-aware feature fusion using a UNet2D-based diffusion model. Experiments on the GoPro dataset demonstrate significant improvements over the baseline Iterative Filter Adaptive Network (IFAN), with Peak Signal-to-Noise Ratio (PSNR) increasing from 8.28 to 34.51 and SSIM from 0.53 to 0.9702, confirming the efficacy of our method for context-aware image deblurring.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages467-471
Number of pages5
ISBN (Electronic)9798331542948
DOIs
Publication statusPublished - 2025
Event2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Kuala Lumpur, Malaysia
Duration: 27 Jun 202528 Jun 2025

Publication series

Name2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings

Conference

Conference2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025
Country/TerritoryMalaysia
CityKuala Lumpur
Period27/06/2528/06/25

Keywords

  • CLIP
  • Image Deblurring
  • Infusion
  • Language
  • Stable Diffusion
  • Vision

Fingerprint

Dive into the research topics of 'Infusing Multimodal Latent Embedding for Image Deblurring'. Together they form a unique fingerprint.

Cite this