Wiki source code of DiffRhythm Musik gen AI

Last modified by René Schmidt on 2025/04/02 14:39

Show last authors
1 [[~[~[image:https://github.com/ASLP-lab/DiffRhythm/raw/main/src/DiffRhythm_logo.jpg~|~|width="400"~]~]>>url:https://github.com/ASLP-lab/DiffRhythm/blob/main/src/DiffRhythm_logo.jpg]]
2
3
4
5 = Di♪♪Rhythm: Blazingly Fast and Embarrassingly Simple
6 End-to-End Full-Length Song Generation with Latent Diffusion =
7
8
9
10 Ziqian Ning, Huakang Chen, Yuepeng Jiang, Chunbo Hao, Guobin Ma, Shuai Wang, Jixun Yao, Lei Xie†
11
12 [[Huggingface Space Demo>>url:https://huggingface.co/spaces/ASLP-lab/DiffRhythm]]
13 📑 [[Paper>>url:https://arxiv.org/abs/2503.01183]] | 📑 [[Demo>>url:https://aslp-lab.github.io/DiffRhythm.github.io/]] | 💬 [[WeChat (微信)>>url:https://github.com/ASLP-lab/DiffRhythm/blob/main/src/contact.md]] 
14
15 DiffRhythm (Chinese: 谛韵, Dì Yùn) is the //**first**// open-sourced diffusion-based music generation model that is capable of creating full-length songs. The name combines "Diff" (referencing its diffusion architecture) with "Rhythm" (highlighting its focus on music and song creation). The Chinese name 谛韵 (Dì Yùn) phonetically mirrors "DiffRhythm", where "谛" (attentive listening) symbolizes auditory perception, and "韵" (melodic charm) represents musicality.
16
17 [[~[~[image:https://github.com/ASLP-lab/DiffRhythm/raw/main/src/diffrhythm.jpg~]~]>>url:https://github.com/ASLP-lab/DiffRhythm/blob/main/src/diffrhythm.jpg]]
18
19
20 == News and Updates ==
21
22
23 * 📌 Join Us on Discord! [[~[~[image:https://camo.githubusercontent.com/de7128d1d75faa616ab515ee182d1ca0e9e66291cac30f837340a71989787d6f/68747470733a2f2f646362616467652e6c696d65732e70696e6b2f6170692f7365727665722f68747470733a2f2f646973636f72642e67672f765544347a6754704a61~|~|alt="Discord"~]~]>>url:https://discord.gg/vUD4zgTpJa]]
24 * (((
25 **2025.3.15 🔥** **DiffRhythm-full Official Release: Complete Music Generation!**
26
27 The wait is over - **285s full-length music generation** is now live!
28
29 //The symphony evolves. What impossible music will you compose next?//
30 )))
31 * **2025.3.11 💻** DiffRhythm can now run on MacOS!
32 * (((
33 **2025.3.9 🔥** **DiffRhythm Update: Text-to-Music and Pure Music Generation!**
34
35 We're excited to announce two groundbreaking features now live in our open-source music model:
36
37 🎯 **Text-Based Style Prompts**
38 Describe styles/scenes in words (e.g., Jazzy Nightclub Vibe, Pop Emotional Piano or Indie folk ballad, coming-of-age themes, acoustic guitar picking with harmonica interludes) — //no audio reference needed!//
39
40 🎧 **Instrumental Mode**
41 Generate pure music with wild prompts like:
42
43 {{{"Arctic research station, theremin auroras dancing with geomagnetic storms" }}}
44 )))
45
46 * (((
47 ✨ Special Thanks to community contributor @Jourdelune for implementing these features via #PR29!
48
49 **Full Release Notes**: See [[src/update_alert.md>>url:https://github.com/ASLP-lab/DiffRhythm/blob/main/src/update_alert.md]] for details, demos, and roadmap.
50
51 Break the rules. Make music that shouldn't exist.
52 )))
53 * **2025.3.7 🔥** **DiffRhythm** is now officially licensed under the **Apache 2.0 License**! 🎉 As the first diffusion-based music generation model, DiffRhythm opens up exciting new possibilities for AI-driven creativity in music. Whether you're a researcher, developer, or music enthusiast, we invite you to explore, innovate, and build upon this foundation.
54 * **2025.3.6 🔥** The local deployment guide is now available.
55 * **2025.3.4 🔥** We released the [[DiffRhythm paper>>url:https://arxiv.org/abs/2503.01183]] and [[Huggingface Space demo>>url:https://huggingface.co/spaces/ASLP-lab/DiffRhythm]].
56
57 == TODOs ==
58
59
60 * Dynamic length control
61 * Vocals only
62 * Song extension
63 * Support Colab.
64 * Gradio support.
65 * Support Docker.
66 * Release DiffRhythm-full.
67 * Release training code.
68 * Support local deployment.
69 * Release paper to Arxiv.
70 * Online serving on Hugging Face Space.
71
72 == Model Versions ==
73
74
75 |=Model|=HuggingFace
76 |DiffRhythm-base (1m35s)|[[https:~~/~~/huggingface.co/ASLP-lab/DiffRhythm-base>>url:https://huggingface.co/ASLP-lab/DiffRhythm-base]]
77 |DiffRhythm-full (4m45s)|[[https:~~/~~/huggingface.co/ASLP-lab/DiffRhythm-full>>url:https://huggingface.co/ASLP-lab/DiffRhythm-full]]
78 |DiffRhythm-vae|[[https:~~/~~/huggingface.co/ASLP-lab/DiffRhythm-vae>>url:https://huggingface.co/ASLP-lab/DiffRhythm-vae]]
79
80 == Docker installation ==
81
82
83 You just need the 3 files inside the folder docker. Do as it follows:
84
85 * Clone the project or copy the files
86 * cd into the folder
87 * Edit your docker compose biding folders
88 * docker compose up -d (or docker-compose up -d depending on your version)
89 * docker exec -it DiffRhythm bash
90
91 You will be in the terminal ready for use. Just go to /home/app/scripts and run infer_prompt_ref.sh
92
93 == Inference ==
94
95
96 Following the steps below to clone the repository and install the environment.
97
98 {{{# clone and enter the repositry
99 git clone https://github.com/ASLP-lab/DiffRhythm.git
100 cd DiffRhythm
101
102 # install the environment
103
104 ## espeak-ng
105 # For Debian-like distribution (e.g. Ubuntu, Mint, etc.)
106 sudo apt-get install espeak-ng
107 # For RedHat-like distribution (e.g. CentOS, Fedora, etc.)
108 sudo yum install espeak-ng
109 # For MacOS
110 brew install espeak-ng
111 # For Windows
112 # Please visit https://github.com/espeak-ng/espeak-ng/releases to download .msi installer
113
114 ## create python environment
115 conda create -n diffrhythm python=3.10
116 conda activate diffrhythm
117
118 ## OR you can use classic Python virtual enviroment instead of conda
119 python -m venv venv
120 # activate venv on Linux
121 source venv/bin/activate
122 # activate venv on Windows
123 venv\Scripts\activate
124
125 ## install requirements
126 pip install -r requirements.txt}}}
127
128 On Linux you can now simply use the inference script:
129
130 {{{# For inference using a reference WAV file
131 bash scripts/infer_wav_ref.sh
132 # For inference using a text prompt reference
133 bash scripts/infer_prompt_ref.sh}}}
134
135 But before running the inference on Windows, make sure you set the user enviroment variables:
136 PHONEMIZER_ESPEAK_LIBRARY -> C:\Program Files\eSpeak NG\libespeak-ng.dll
137 PHONEMIZER_ESPEAK_PATH -> C:\Program Files\eSpeak NG
138 Change C:\Program Files\eSpeak NG to your eSpeak installation directory and reboot your PC to apply changes.
139
140 //Installing Japanese voices, mbrola binaries and unpacking an mbrola_ph folder (as described [[here>>url:https://github.com/ASLP-lab/DiffRhythm/issues/15]] and [[here>>url:https://github.com/ASLP-lab/DiffRhythm/issues/22]]) are **no longer required** when running on Windows. See [[#17 (comment)>>url:https://github.com/ASLP-lab/DiffRhythm/issues/17#issuecomment-2705058729]], [[this>>url:https://github.com/ASLP-lab/DiffRhythm/commit/2ea9424274df10670ddc613b5d61cc16d13e2b88]] and [[this commit>>url:https://github.com/ASLP-lab/DiffRhythm/commit/1ad7229e1a774c9a2a0c4888103dd4ea7176aebb]].//
141
142 After this, you will also be able to run inference scripts on Windows (please note that English lyrics will be used here):
143
144 {{{rem : For inference using a reference WAV file
145 call scripts\infer_wav_ref.bat
146 rem : For inference using a text prompt reference
147 call scripts\infer_prompt_ref.bat}}}
148
149 Example files of lrc and reference audio can be found in infer/example.
150
151 You can use [[the tools>>url:https://huggingface.co/spaces/ASLP-lab/DiffRhythm]] we provide on huggingface to generate the lrc.
152
153 **Note that DiffRhythm-base requires a minimum of 8G of VRAM. To meet the 8G VRAM requirement, use the ~-~-chunked argument when running the inference. Higher VRAM may be required if chunked decoding is disabled.**
154
155 == Training ==
156
157
158 Coming soon...
159
160 == License & Disclaimer ==
161
162
163 DiffRhythm (code and DiT weights) is released under the [[Apache License 2.0>>url:https://www.apache.org/licenses/LICENSE-2.0]]. This open-source license allows you to freely use, modify, and distribute the model, as long as you include the appropriate copyright notice and disclaimer.
164
165 We do not make any profit from this model. Our goal is to provide a high-quality base model for music generation, fostering innovation in AI music and contributing to the advancement of human creativity. We hope that DiffRhythm will serve as a foundation for further research and development in the field of AI-generated music.
166
167 DiffRhythm enables the creation of original music across diverse genres, supporting applications in artistic creation, education, and entertainment. While designed for positive use cases, potential risks include unintentional copyright infringement through stylistic similarities, inappropriate blending of cultural musical elements, and misuse for generating harmful content. To ensure responsible deployment, users must implement verification mechanisms to confirm musical originality, disclose AI involvement in generated works, and obtain permissions when adapting protected styles.