1.湘江实验室,湖南 长沙 410205
2.湖南工商大学人工智能与先进计算学院,湖南 长沙 410205
3.湖南工商大学智能工程与智能制造学院,湖南 长沙 410205
[ "姜林(1977- ),男,博士,湖南工商大学人工智能与先进计算学院教授,主要研究方向为智能语音处理、机器人应用。" ]
[ "苗向阳(1998- ),男,湖南工商大学智能工程与智能制造学院硕士生,主要研究方向为智能语音处理、音效生成。" ]
[ "洪紫荆(1997- ),女,湖南工商大学人工智能与先进计算学院硕士生,主要研究方向为智能语音处理、音效生成。" ]
收稿:2025-03-23,
修回:2025-05-26,
纸质出版:2025-12-15
移动端阅览
姜林,苗向阳,洪紫荆.仅用单一短音频训练的音效生成GAN模型[J].智能科学与技术学报,2025,07(04):468-483.
JIANG Lin,MIAO Xiangyang,HONG Zijing.A GAN generation method of sound effect using only a single short audio for training[J].Chinese Journal of Intelligent Science and Technology,2025,07(04):468-483.
姜林,苗向阳,洪紫荆.仅用单一短音频训练的音效生成GAN模型[J].智能科学与技术学报,2025,07(04):468-483. DOI: 10.11959/j.issn.2096-6652.202524.
JIANG Lin,MIAO Xiangyang,HONG Zijing.A GAN generation method of sound effect using only a single short audio for training[J].Chinese Journal of Intelligent Science and Technology,2025,07(04):468-483. DOI: 10.11959/j.issn.2096-6652.202524.
针对音效生成中音频逼真度低、风格多样性欠缺的问题,提出了一种基于多频带注意力机制的生成对抗网络(generative adversarial network,GAN)模型。首先,采用多频带扩展模式提取不同采样率的音频特征,并引入RaHingeGAN(relativistic average hinge GAN)损失函数提高音频生成稳定性。其次,结合Transformer注意力机制增强谐波信息和频谱结构表达,并引入Alpha Dropout自适应正则层缓解过拟合。最后,设计音频风格迁移模块以增强风格可控性,在特征学习过程中,融合梅尔频率倒谱系数(Mel frequency cepstral coefficient,MFCC)、伽玛通频率倒谱系数(Gammatone frequency cepstrum coefficient,GFCC)及其多阶差分系数捕捉音频信号动态特征。实验表明,基于多频带注意力机制的GAN模型在音效逼真度、风格多样性和生成稳定性方面均优于现有模型,可有效提升音效生成质量。
To address the issues of low audio realism and insufficient style diversity in sound effect generation
a generative adversarial network (GAN) model based on a multi-band attention mechanism was proposed. Firstly
a multi-band expansion mode was adopted to extract audio features at different sampling rates
and a relativistic average hinge GAN (RaHingeGAN) loss function was introduced to improve audio generation stability. Secondly
a Transformer attention mechanism was incorporated to enhance the expression of harmonic information and spectral structure
while an Alpha Dropout adaptive regularization layer was applied to mitigate overfitting. Finally
an audio style transfer module was designed to enhance style controllability. During the feature learning process
Mel frequency cepstral coefficient (MFCC)
Gammatone frequency cepstrum coefficient (GFCC)
and their higher-order differential coefficients were fused to capture the dynamic characteristics of the audio signal. Experiments demonstrated that the proposed GAN model based on the multi-band attention mechanism outperformed existing models in terms of sound effect realism
style diversity
and generation stability
effectively improving the quality of sound effect generation.
NATSIOU A, O'LEARY S. Audio representations for deep learning in sound synthesis: a review[C]//Proceedings of the 2021 IEEE/ACS 18th International Conference on Computer Systems and Applications (AICCSA). Piscataway: IEEE Press, 2021: 1-8.
黄峻, 林飞, 杨静, 等. 生成式AI的大模型提示工程: 方法、现状与展望[J]. 智能科学与技术学报, 2024, 6(2): 115-133.
HUANG J, LIN F, YANG J, et al. From prompt engineering to generative artificial intelligence for large models: the state of the art and perspective[J]. Chinese Journal of Intelligent Science and Technology, 2024, 6(2): 115-133.
任维俊. AI音频生成在广播电视领域的应用研究[J]. 电声技术, 2024, 48(10): 110-112.
REN W J. Research on the application of AI audio generation in the field of radio and television[J]. Audio Engineering, 2024, 48(10): 110-112.
林义超, 张泽. 非渲染技术语境下沉浸式音频空间构建方案的探索与思考[J]. 现代电影技术, 2024(12): 42-49.
LIN Y C, ZHANG Z. An exploration and reflection on immersive audio space configuration schemes in non-rendering technology context[J]. Advanced Motion Picture Technology, 2024(12): 42-49.
ZHAO Y J, XIA X J, TOGNERI R. Applications of deep learning to audio generation[J]. IEEE Circuits and Systems Magazine, 2019, 19(4): 19-38.
倪清桦, 鲁越, 林飞, 等. 平行音乐: 大模型时代的人机混合音乐创演[J]. 智能科学与技术学报, 2024, 6(2): 150-163.
NI Q H, LU Y, LIN F, et al. Parallel music: human-machine hybrid music creation and performance in the era of large models[J]. Chinese Journal of Intelligent Science and Technology, 2024, 6(2): 150-163.
王健宗, 张旭龙, 姜桂林, 等. 基于分层联邦框架的音频模型生成技术研究[J]. 智能系统学报, 2024, 19(5): 1331-1339.
WANG J Z, ZHANG X L, JIANG G L, et al. Research on audio model generation technology based on a hierarchical federated framework[J]. CAAI Transactions on Intelligent Systems, 2024, 19(5): 1331-1339.
靳聪, 王洁, 郭子淳, 等. 融合音画同步的唇形合成研究[J]. 智能科学与技术学报, 2023, 5(3): 397-405.
JIN C, WANG J, GUO Z C, et al. Lipsynthesis incorporating audio-visual synchronisation[J]. Chinese Journal of Intelligent Science and Technology, 2023, 5(3): 397-405.
GOEL K, GU A, DONAHUE C, et al. It's raw! audio generation with state-space models[C]//Proceedings of the 39th International Conference on Machine Learning. New York: PMLR, 2022: 7616-7633.
GU A, GOEL K, RÉ C. Efficiently modeling long sequences with structured state spaces[J]. arXiv preprint, 2021, arXiv: 2111.00396.
YANG D C, YU J W, WANG H L, et al. Diffsound: discrete diffusion model for text-to-sound generation[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, 31: 1720-1733.
ZHOU Y P, WANG Z W, FANG C, et al. Visual to sound: generating natural sound for videos in the wild[C]//Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 3550-3558.
谢志峰, 孙络祎, 孙郁洲, 等. 时序对齐视觉特征映射的音效生成方法[J]. 计算机辅助设计与图形学学报, 2022, 34(10): 1506-1514.
XIE Z F, SUN L Y, SUN Y Z, et al. Sound generation method with timing-aligned visual feature mapping[J]. Journal of Computer-Aided Design & Computer Graphics, 2022, 34(10): 1506-1514.
XIE Z Y, LI B H, XU X N, et al. Enhancing audio generation diversity with visual information[C]//Proceedings of the ICASSP 2024—2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2024: 866-870.
LIU X B, IQBAL T, ZHAO J Z, et al. Conditional sound generation using neural discrete time-frequency representation learning[C]//Proceedings of the 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP). Piscataway: IEEE Press, 2021: 1-6.
KREUK F, SYNNAEVE G, POLYAK A, et al. AudioGen: textually guided audio generation[J]. arXiv preprint, 2022, arXiv: 2209.15352.
SHEFFER R, ADI Y. I hear your true colors: image guided audio generation[C]//Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2023: 1-5.
BORSOS Z, MARINIER R, VINCENT D, et al. AudioLM: a language modeling approach to audio generation[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, 31: 2523-2533.
GUO Z F, MAO J G, TAO R, et al. Audio generation with multiple conditional diffusion model[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(16): 18153-18161.
LIU H H, CHEN Z H, YUAN Y, et al. AudioLDM: text-to-audio generation with latent diffusion models[J]. arXiv preprint, 2023: arXiv: 2301.12503.
HUANG R J, HUANG J W, YANG D C, et al. Make-an-audio: text-to-audio generation with prompt-enhanced diffusion models[C]//Proceedings of the 40th International Conference on Machine Learning. New York: PMLR, 2023: 13916-13932.
AGOSTINELLI A, DENK T I, BORSOS Z, et al. MusicLM: generating music from text[J]. arXiv preprint, 2023: arXiv: 2301.11325.
LIU S G, LI S J, CHENG H N. Towards an end-to-end visual-to-raw-audio generation with GAN[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(3): 1299-1312.
GHOSE S, PREVOST J J. FoleyGAN: visually guided generative adversarial network-based synchronous sound generation in silent videos[J]. IEEE Transactions on Multimedia, 2022, 25: 4508-4519.
BAOUEB T, LIU H C, FONTAINE M, et al. SpecDiff-GAN: a spectrally-shaped noise diffusion GAN for speech and music synthesis[C]//Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2024: 986-990.
GRESHLER G, SHAHAM T R, MICHAELI T. Catch-A-waveform: learning to generate audio from a single short example[C]//Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS 2021). New York: Curran Associates Inc., 2021: 1-13.
WU S L, YANG Y H. MuseMorphose: full-song and fine-grained piano music style transfer with one transformer VAE[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, 31: 1953-1967.
KOO J, MARTÍNEZ-RAMÍREZ M A, LIAO W H, et al. Music mixing style transfer: a contrastive learning approach to disentangle audio effects[C]//Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2023: 1-5.
ZHANG Y, HUANG R J, LI R Q, et al. StyleSinger: style transfer for out-of-domain singing voice synthesis[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(17): 19597-19605.
WU Y X, HE Y F, LIU X L, et al. Transplayer: timbre style transfer with flexible timbre control[C]//Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2023: 1-5.
SHAHAM T R, DEKEL T, MICHAELI T. SinGAN: learning a generative model from a single natural image[C]//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE Press, 2019: 4570-4580.
VASWANI A , SHAZEER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. New York: Curran Associates Inc.,2017: 6000-6010.
李牧, 杨宇恒, 柯熙政. 基于混合特征提取与跨模态特征预测融合的情感识别模型[J]. 计算机应用, 2024, 44(1): 86-93.
LI M, YANG Y H, KE X Z. Emotion recognition model based on hybrid-mel gama frequency cross-attention transformer modal[J]. Journal of Computer Applications, 2024, 44(1): 86-93.
WEBBER J J, VALENTINI-BOTINHAO C, WILLIAMS E, et al. Autovocoder: fast waveform generation from a learned speech representation using differentiable digital signal processing[C]//Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2023: 1-5.
FONT F, ROMA G, SERRA X. Freesound technical demo[C]//Proceedings of the 21st ACM International Conference on Multimedia. New York: ACM, 2013: 411-412.
LOSHCHILOV I, HUTTER F. Decoupled weight decay regularization[J]. arXiv preprint, 2017, arXiv: 1711.05101.
BARAHONA-RÍOS A, COLLINS T. SpecSinGAN: sound effect variation synthesis using single-image GANs[J]. arXiv preprint, 2021, arXiv: 2110.07311.
GRAY K L H, HAFFEY A, MIHAYLOVA H L, et al. Lack of privileged access to awareness for rewarding social scenes in autism spectrum disorder[J]. Journal of Autism and Developmental Disorders, 2018, 48(10): 3311-3318.
ROUDER J N, SPECKMAN P L, SUN D C, et al. Bayesian t tests for accepting and rejecting the null hypothesis[J]. Psychonomic Bulletin & Review, 2009, 16(2): 225-237.
0
浏览量
144
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621
