YuE2: Open-Source AI Song Generator With an Editable Score Rivals Suno v5

YuE2 song generation AI

A team from HKUST, HKGAI and the Multimodal Art Projection (M-A-P) community has released YuE2, an open-source AI song generator. It first composes a song as a readable score and then performs it as a finished track with vocals. You can edit the score by hand or through an LLM agent to get a new version of the song. If you transcribe the notes of someone else’s recording, the model will perform a cover in a different style without any fine-tuning.

In a blind listening test, experts preferred YuE2 in best-of-8 mode (the best of eight candidates) over Suno v4.5 in 57.3% of comparisons versus 30.5%. Against Suno v5 it was almost a tie: 40.4% versus 39.9%. On the WildSongBench benchmark, this mode scores 6.96 on SongBench, more than any of the 15 other models in the comparison. However, the authors do not consider its lead over the closest rivals, Mureka 9 (6.94) and Suno v5 (6.87), statistically significant. The tests also show that the score makes songs better: with it, experts chose the track in 49.3% of comparisons, without it in 34.6%.

The weights of YuE2 (3.58B parameters) and of the auxiliary MERT2, SheetSage2 and VAE models are available on Hugging Face under the non-commercial CC BY-NC 4.0 license. At the same time, the authors separately allow individual creators and musicians to sell the songs they generate. They also released the benchmark together with its evaluation code. The GitHub repository hosts the inference code and the yue2-music skill for LLM agents, while the code of the first YuE now lives in a separate branch. Song samples are available on the project page.

Song quality index vs text alignment index on WildSongBench: YuE2 and YuE2 (Bo8) next to Suno v5, Suno v6 and Mureka 9, with open models lower and to the left
In song quality and prompt adherence, YuE2 sits closer to Suno and Mureka than to other open models

Why an AI song generator needs a score

You can describe a song at several levels: lyrics and mood, notes and chords, performance and timbre, and finally the waveform. The authors call top-down generation through these levels progressive musical commitment. Symbolic models such as Music Transformer stop at the notes. Audio models such as MusicLM, MusicGen and the first YuE output a recording right away and keep the composition hidden inside. ByteDance’s Seed-Music described a similar path from notes to song, but its implementation is closed. YuE2 goes through all the levels in a single open model and keeps the score available for editing.

Pyramid of generation levels: text and lyrics, symbolic score, semantic tokens, acoustic latents, waveform
Progressive musical commitment: from lyrics to waveform

YuE2 writes the score in ABC notation, which is plain text: tempo, meter, key, bars, verse and chorus labels, chords and notes. The same BPE tokenizer that handles the lyrics also splits ABC into tokens, so both a person and an external language model can edit it.

How YuE2 works

Generation runs in three steps: the score, then semantic tokens, then continuous acoustic latent representations (latents). For these, the model splits audio into 40 ms frames and describes each frame with 64 numbers. Semantic tokens add what is awkward to write down in notes, and a separate variational autoencoder (VAE) turns the latents into 48 kHz stereo audio.

A single 28-layer transformer built as a Mixture-of-Transformers (MoT) handles all three steps: each layer has two experts with their own weights and a shared attention mechanism. The autoregressive (AR) expert writes the score and tokens one at a time, the way a language model writes text. The non-autoregressive (NAR) expert processes all frames at once and gradually turns noise into latents with flow matching. The attention mask is hybrid. The model writes the score and semantic tokens looking only at what it has already generated, while audio frames see all tokens and each other.

Architecture: an AR expert with causal attention writes the ABC score and semantic tokens, a NAR expert with bidirectional attention predicts acoustic latents for the VAE, with offline targets from SheetSage2, MERT2 and the VAE encoder at the bottom
One checkpoint creates songs, edits them and makes covers; only the source of the score changes

The authors trained the model on 346,000 hours of music, mostly CC0-licensed recordings and synthetic data. Training mixed four tasks, from the full chain down to audio generation without a score or tokens, so a single checkpoint can work with or without a plan.

Where the training scores come from: SheetSage2 and MERT2

Ordinary recordings carry no annotations of notes and chords, so the authors built two tools. SheetSage2 is an audio-to-sheet-music transcriber: it listens to the whole song and outputs the melody, chords, key, beat grid and structure in ABC. Among the compared transcribers, it is the best on 12 of 15 benchmark–metric pairs. For example, vocal melody F1 on RWC-Pop rose from 62.71 in the first version to 82.51. The authors trained the MERT2 encoder (632M parameters) on 700,000 hours of audio to predict codes of masked fragments derived from two frozen encoders, MuQ and Qwen2-Audio. As a result, it became the best on 14 of 15 metrics of the MARBLE benchmark. Its quantized branch produces the semantic tokens for YuE2 at 375 bits per second.

YuE2 vs Suno and Mureka: closed-model quality from an open model

The comparison used WildSongBench: 192 real user requests in Chinese and English. Automatic metrics measured song quality and musicality (SongBench, SongEval), production quality (AudioBox), prompt adherence and lyric intelligibility (PER). YuE2 appears twice in the table, but it is the same model. In the YuE2 row it generates two candidates per request, like all competitors, and the one with clearer lyrics counts. In the YuE2 (Bo8) row there are eight candidates (best-of-8), and SongBench Musicality picks the best one. With two candidates, YuE2 leads the open models on SongBench (6.73), even though six of its eight open competitors have more parameters. On AudioBox PQ, which played no role in selection, both modes beat every closed model.

Table: SongBench, SongEval, AudioBox PQ, MuLan, AllMusicCaps, Q3O and PER for Suno v4.5, v5, v5.5, v6, v6 Wild, MiniMax Music 2.6, Mureka 9, YuE2 and YuE2 (Bo8)
YuE2 (Bo8) leads on SongBench and Q3O, Suno v5 on MuLan and AllMusicCaps

The authors treat blind listening as the main evidence of quality. In the test, 62 paid experts with conservatory training or audio AI research experience gave 3,602 pairwise ratings, 2,545 of them in comparisons with closed models. Against Mureka 9 the result is also nearly even. Experts did prefer Suno v6, though: 59.3% versus 31.6%, even though YuE2 scores higher on SongBench. For audio quality, experts chose YuE2 (Bo8) on average in 58.9% of comparisons with six closed models (ties split equally). Overall, YuE2 looks like an open-source Suno alternative on par with Suno v5, but not with its latest versions.

Expert preference shares for overall quality and audio quality: YuE2 and YuE2 best-of-8 vs Suno v4.5, v5, v5.5, v6, v6 Wild and Mureka 9
YuE2 (Bo8) beats Suno v4.5, nearly ties Suno v5 and Mureka 9, and trails newer Suno versions

What the score and the unified architecture add

To test the idea itself, the authors ran the same checkpoint with and without a plan on identical prompts. Experts preferred the song with a score in 49.3% of responses and the one without it in 34.6% (p = 0.007). For melody the split was 44.0% versus 29.5%, and for chord progression 39.0% versus 21.0%.

Expert preferences for generation with and without a score: overall quality, musicality, melody and chord progression
With a score, experts chose the song more often on all four criteria

An example shows where the difference comes from: with a plan, the verse develops a single motif and the chorus brings it back an octave higher. Without a plan, the verse repeats the same B–D–E–E figure four times.

Verse and chorus score excerpts of two songs with the same prompt: with symbolic planning the motif develops and returns in the chorus, without planning the phrase repeats
With a plan (left), the verse motif returns in the chorus an octave higher; without a plan, the phrase simply repeats

The second test concerns the architecture. Many song generators consist of two parts: a language model (LM) for tokens and a diffusion transformer (DiT) for audio. The authors trained such a pair, 1.7B parameters each, on the same amount of data and with the same planning. Experts chose the unified MoT in 53.4% of responses versus 35.6% (p = 0.0084).

Score editing, AI covers and an LLM agent

A score is only useful if the model follows it. The authors ran generated songs through SheetSage2. Melody similarity to the plan is 0.9464 and chord similarity 0.9246. For a song rendered from a different score with the same prompt, the figures drop to 0.2160 and 0.2473. Edits stay local. After you replace the chords in the chorus, the recording matches the new chords by 79.54%, while the melody still matches the original by 93.43%. The SongBench score barely changes after edits.

For an AI cover, SheetSage2 transcribes a score from a recording and YuE2 performs it in a new style. The authors tested this on 948 songs from SHS100K that were absent from the training data. From a generated version, the CLEWS cover-song retrieval model found other performances of the same song with an mAP of 0.647. For SongEcho, a dedicated cover model, this figure is 0.419. Without chords in the score, recognizability drops to 0.598, and without a score almost to zero (0.006), although musicality rises.

Table: retrieving the source song from AI covers with CLEWS and Discogs-VINet, plus style alignment and quality for SongEcho, ACE-Step 1.5 and YuE2 with the full score, without chords and without a score
Without chords, recognizability drops to 0.598 and without a score to 0.006, while style match and audio quality rise

A readable score also suits LLM agents. In a demo, the song The Last Train went from Mandarin pop to English-language jazz over 14 versions. The agent edited the notes, style and lyrics at the researcher’s request, and YuE2 performed each version.

On lyric intelligibility (PER 0.0844), YuE2 trails seven of its 15 competitors, including Suno v4.5 (0.0580) and the best open model, MiniMax Music 3 (0.0627). The 6.96 record requires eight generations per song and partly reflects the fact that SongBench Musicality picks the best candidate.

How to try YuE2 for free and run it locally

You can create a song with YuE2 without installing anything in the free web demo from NOIZ, one of the project’s partners. The number of free generations there is limited, though. The model also runs in third-party Hugging Face Spaces, for example YuE2-3B Music Generator. The model card lists English and Chinese as the supported languages.

For local use, the authors specify Linux, Python 3.10+, a 24 GB NVIDIA GPU with BF16 support and 24 GB of RAM. In tests, however, peak VRAM was about 11 GiB, rising to 14.1 GiB at the maximum context length. On an RTX 4090, a 3.6-minute song takes 71 seconds to generate. The yue2_infer package installs via pip from the GitHub repository or as a prebuilt wheel from Hugging Face. It takes a style description and lyrics tagged with verses and choruses, so it works as a lyrics-to-song generator. The cot parameter switches between the full plan (full), a melody-only plan for covers (melody) and generation without a score (off). For editing or covers, you pass a ready-made ABC score. ComfyUI has native YuE2 nodes and an official text-to-music workflow.

Can you use YuE2 songs commercially?

The authors distribute the YuE2-3B and YuE2-Vae weights under the CC BY-NC 4.0 license, which generally prohibits commercial use. However, on September 16 the authors added a separate permission for individual users, content creators and musicians. They can generate songs for free and publish, sell and license them, including on commercial platforms, without paying royalties to the developers. The condition is responsible use: users must not apply the model or the generated songs to illegal or deceptive purposes such as fraud or impersonating another person. The permission does not cover companies, which need a separate commercial license from the authors to use the weights. No one may sell or commercially redistribute the weights themselves.


bnr2mob
Subscribe
Notify of
guest

0 Comments
Oldest
Newest Most Voted