WaveConvNeXt: Wavelet-Guided Multi-Scale Feature Extraction with Frequency-Domain Attention for High-Fidelity Image Classification and Reconstruction
Keywords:
Deep learning; Image classification; Wavelet transform; Frequency-domain attention; ConvNeXt; Vision Transformer; Image super-resolution; Multi-scale feature extraction; ImageNetAbstract
Deep convolutional neural networks and Vision Transformers have achieved remarkable accuracy on large-scale image classification benchmarks, yet they predominantly operate in the spatial domain and fail to exploit the rich frequency-domain structure of natural images. Frequency information — encoded by wavelets as localized time-frequency representations — captures multi-resolution edge, texture, and periodic patterns that are complementary to the spatial receptive fields of convolutions and the global attention of Transformers. In this paper, we propose WaveConvNeXt, a hybrid deep learning architecture that integrates discrete wavelet transform (DWT) decompositions directly into the feature extraction pipeline of ConvNeXt-V2, augmented by a Frequency-Domain Attention (FDA) module that gates feature maps according to their energy distribution across wavelet sub-bands. WaveConvNeXt decomposes each feature map into four DWT sub-bands (LL, LH, HL, HH) at each stage, applies sub-band-specific convolutions to extract directional frequency features, and recombines them using learned FDA weights that adapt to the spatial frequency content of each input. We additionally propose a Wavelet Image Super-Resolution (WISR) decoder that leverages the preserved high-frequency sub-bands to reconstruct high-fidelity upsampled images as an auxiliary training signal. We evaluate WaveConvNeXt on three large-scale benchmarks: ImageNet-1K (1.28 million images, 1,000 classes), CIFAR-100 (60,000 images, 100 classes), and the DIV2K image super-resolution dataset (800 high-resolution training images). On ImageNet-1K, WaveConvNeXt-Base achieves 84.7% top-1 accuracy, surpassing ConvNeXt-V2-Base (83.7%) by 1.0% and Swin-Transformer-Base (83.5%) by 1.2%, with only 5.3% additional parameters. On CIFAR-100, WaveConvNeXt achieves 91.8% top-1 accuracy. On DIV2K ×4 super-resolution, WISR achieves PSNR of 33.24 dB and SSIM of 0.913, outperforming DRN and SwinIR baselines. Extensive ablation studies confirm that frequency-domain attention contributes +1.4% on ImageNet and +0.8% on CIFAR-100 over the ConvNeXt-V2 backbone alone.
References
[1] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Commun. ACM, vol. 60, no. 6, pp. 84–90, 2017.
[2] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Proc. CVPR, pp. 248–255, 2009.
[3] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. ICLR, 2015.
[4] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. CVPR, pp. 770–778, 2016.
[5] M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. ICML, pp. 6105–6114, 2019.
[6] A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021.
[7] Z. Liu et al., “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proc. ICCV, pp. 10012–10022, 2021.
[8] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A ConvNet for the 2020s,” in Proc. CVPR, pp. 11976–11986, 2022.
[9] S. Woo et al., “ConvNeXt V2: Co-designing and scaling ConvNets with masked autoencoders,” in Proc. CVPR, 2023.
[10] S. G. Mallat, “A theory for multiresolution signal decomposition: The wavelet representation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 11, no. 7, pp. 674–693, 1989.
[11] I. Daubechies, “Orthonormal bases of compactly supported wavelets,” Commun. Pure Appl. Math., vol. 41, no. 7, pp. 909–996, 1988.
[12] L. Theis, W. Shi, A. Cunningham, and F. Huszár, “Lossy image compression with compressive autoencoders,” in Proc. ICLR, 2017.
[13] D. Nie et al., “Medical image synthesis with deep convolutional adversarial networks,” IEEE Trans. Biomed. Eng., vol. 65, no. 12, pp. 2720–2730, 2018.
[14] W. Qian, D. Dong, Y. Li, and J. You, “Deep learning in steganography and steganalysis from 2015 to 2018,” arXiv:1904.01444, 2019.
[15] M. Tan and Q. V. Le, “EfficientNetV2: Smaller models and faster training,” in Proc. ICML, pp. 10096–10106, 2021.
[16] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Proc. ICML, pp. 10347–10357, 2021.
[17] Z. Tu et al., “MaxViT: Multi-axis vision transformer,” in Proc. ECCV, pp. 459–479, 2022.
[18] Q. Li, L. Shen, S. Guo, and Z. Lai, “Wavelet integrated CNNs for noise-robust image classification,” in Proc. CVPR, pp. 7245–7254, 2020.
[19] Y. Li, Z. Zhang, J. Tan, and W. Chen, “Wavelet-based feature pyramid for medical image segmentation,” Pattern Recognit. Lett., vol. 140, pp. 1–8, 2020.
[20] X. Liu, M. Tanaka, and M. Okutomi, “Noise level estimation using weak textured patches of a single noisy image,” in Proc. ICIP, pp. 665–668, 2012.
[21] Q. Qin et al., “FcaNet: Frequency channel attention networks,” in Proc. ICCV, pp. 783–792, 2021.
[22] L. Gueguen, A. Sergeev, B. Kadlec, R. Liu, and J. Yosinski, “Faster neural networks straight from JPEG,” in Proc. NeurIPS, pp. 3937–3948, 2018.
[23] C. Dong, C. C. Loy, K. He, and X. Tang, “Learning a deep convolutional network for image super-resolution,” in Proc. ECCV, pp. 184–199, 2014.
[24] B. Lim, S. Son, H. Kim, S. Nah, and K. Mu Lee, “Enhanced deep residual networks for single image super-resolution,” in Proc. CVPR Workshops, pp. 136–144, 2017.
[25] Y. Guo, J. Chen, J. Wang, Q. Chen, J. Cao, Z. Deng, Y. Xu, and M. Tan, “Closed-loop matters: Dual regression networks for single image super-resolution,” in Proc. CVPR, pp. 5407–5416, 2020.
[26] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “SwinIR: Image restoration using Swin transformer,” in Proc. ICCV Workshops, pp. 1833–1844, 2021.
[27] X. Wang et al., “Real-ESRGAN: Training real-world blind super-resolution with pure synthetic data,” in Proc. ICCV Workshops, pp. 1905–1914, 2021.
[28] E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le, “AutoAugment: Learning augmentation strategies from data,” in Proc. CVPR, pp. 113–123, 2019.
[29] H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “Mixup: Beyond empirical risk minimization,” in Proc. ICLR, 2018.
[30] S. Yun, D. Han, S. Chun, S. Oh, Y. Yoo, and J. Choe, “CutMix: Regularization strategy to train strong classifiers with localizable features,” in Proc. ICCV, pp. 6023–6032, 2019.
[31] A. Krizhevsky, “Learning multiple layers of features from tiny images,” Tech. Rep., Univ. of Toronto, 2009.
[32] E. Agustsson and R. Timofte, “NTIRE 2017 challenge on single image super-resolution: Dataset and study,” in Proc. CVPR Workshops, pp. 126–135, 2017.
[33] P. Martin, P. Roelants, and D. Forsyth, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in Proc. ICCV, vol. 2, pp. 416–423, 2001.
[34] J.-B. Huang, A. Singh, and N. Ahuja, “Single image super-resolution from transformed self-exemplars,” in Proc. CVPR, pp. 5197–5206, 2015.
[35] W. Shi et al., “Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,” in Proc. CVPR, pp. 1874–1883, 2016.
[36] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR, 2019.
[37] H. Touvron, P. Bojanowski, M. Caron, M. Douze, A. Joulin, A. Synnaeve, G. Izacard, and H. Jégou, “ReScaling Vision Transformers: Revisiting existing vision transformers at scale,” in Proc. ICLR, 2022.
[38] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. CVPR, pp. 7132–7141, 2018.