lab

YuE2-3B on Victoria: architecture, inference path and RTX 3050 feasibility

Analysed 2026-09-14 against the released m-a-p/YuE2-3B checkpoint (model repo created 2026-09-09, revision 1a96eca688d6ae5d7f0feb88573fec89920fcd19) and the shipped inference wheel yue2_infer-0.1.5-py3-none-any.whl. The target is Victoria, an NVIDIA GeForce RTX 3050 6GB Laptop GPU (GA107, sm_86, 6144 MiB VRAM, 2560 CUDA cores at 1.24–1.49 GHz, 96-bit bus at 12 Gbps = 144 GB/s memory bandwidth). Victoria was offline for the whole analysis (Tailscale last saw it 15 h earlier; ssh victoria timed out on the first attempt and was not retried), so every latency figure in this report is analytical — derived from the released weights, the released runtime source, and the model card's published RTX 4090 / H800 measurements. Nothing here is a measured 3050 number. The section "What to measure next" contains the script that replaces the estimates with measurements once Victoria is reachable.

Bottom line: YuE2-3B cannot run on Victoria, and the blocker is memory, not speed. The BF16 MoT backbone alone is 6.76 GiB of weights. The pipeline refuses to allocate more than total_memory - 2 GiB, which is 4.00 GiB on a 6 GB card and 6.00 GiB even on an 8 GB card — so the checkpoint does not fit on any RTX 3050 variant, at any setting the released runtime supports. The one built-in compression path (FP8) requires compute capability ≥ 8.9 and is rejected on sm_86. Every other route to fitting requires a different runtime (GGUF/audio.cpp), and the arithmetic for that route says a 3.6-minute song would still take roughly 8.5–15 minutes to render. The 96 GB RTX PRO 6000 already in the benchmark pool is the host that actually fits.

Reproducibility

No GPU was available, so the analysis is a static one over released artifacts. Three sources were read directly rather than inferred:

Model card: https://huggingface.co/m-a-p/YuE2-3B. Code: https://github.com/multimodal-art-projection/YuE2.

Architecture

YuE2-3B is a Mixture-of-Transformers (MoT): one 28-layer backbone in which every layer carries two parallel branches — an autoregressive branch (self_attn, mlp) and a non-autoregressive branch (nar_self_attn, nar_mlp) — plus a separate Oobleck VAE decoder. The AR branch writes the song as text and discrete tokens; the NAR branch runs flow matching to turn those tokens into continuous latents; the VAE turns latents into stereo audio.

property value source
total parameters 3,630,684,224 (3.631 B) summed from the safetensors header
checkpoint size, BF16 7,261,368,448 B = 7.26 GB = 6.763 GiB header / weights_manifest.json
hidden size 2048 config.json
layers 28 config.json
attention heads / KV heads 16 / 8 (GQA 2:1), head_dim 128 config.json
intermediate size 6144 (SwiGLU) config.json
vocabulary 184,704 (Qwen tiktoken base + 32,768 codec tokens) config.json, protocol.py
max positions 24,576 config.json
AR branch weights 1,409.4 M header (sums over layers.*.self_attn, layers.*.mlp)
NAR branch weights 1,409.4 M header (sums over layers.*.nar_*)
embeddings 378.3 M (embed_tokens) header
LM head 378.3 M, untied header (tie_word_embeddings: false)
latent position table 50.3 M (24576 × 2048) header
timestep embedder + vae2llm/llm2vae 4.9 M header

The two branches are exactly equal in size (1,409.4 M each), which is worth knowing: the NAR half of the checkpoint is dead weight during the AR phase and vice versa. No released runtime exploits that yet; nar.py's optional offload_ar flag is the only nod to it, and it swaps branch weights through host RAM rather than keeping only one resident.

The layer code contains a combined branch (modeling_yue2.py:256-295) that computes both branches and merges them with torch.where(mask_3d, ar_out, nar_out). That path is not used at inference — sampling.py takes the AR-only else branch when ar_mask is None, and nar.py:161-167 drives the NAR modules directly, bypassing DecoderLayer.forward entirely. So there is no 2× waste in the released pipeline, but any reimplementation that calls the layer naively will pay it.

The audio tokeniser

The VAE config fixes the latent geometry: downsampling_ratio: 1920 at 48 kHz stereo, i.e. one latent frame = 40.0 ms of audio, 25 frames per second.

quantity value
latent frame rate 25 Hz (40 ms/frame)
latent channels 64
maximum latent frames 24,576 → 983 s ≈ 16.4 min
decoder architecture Oobleck, c_mults=[1,2,4,8,16,32], strides=[2,2,4,4,5,6], channels 64, SnakeBeta, final_tanh: false
decoder weights, FP32 530 MB (m-a-p/YuE2-Vae/model.safetensors)
tiled decode decode_core_frames 1024, halo 16 (512 when memory_budget_gib <= 12)

The VAE is derived from stable-audio-tools (a6ae0cdf, MIT) and NVIDIA's SnakeBeta — the wheel ships both licences. It always runs in FP32; pipeline.py:371 pins "vae_dtype": "float32".

Verified inference path

Four stages, exposed separately as plan() → generate_semantic() → synthesize() → decode(). One pipe(...) call chains all four.

  1. Prompt assembly — protocol.py:108-126. SongRequest.text() concatenates a mode instruction, [Tags], the style string, [Lyrics] and the lyrics. token_prefixes() prepends EOD (151643) and appends ABC_START/ABC_END/MUSIC_START (151847/151848/151851).
  2. Symbolic planning (AR) — plan(), phase "abc". The AR branch generates an ABC melody-and-chord score. Defaults temperature 0.7, top_p 0.9, top_k 30, max_tokens 4096. Skipped entirely when cot="off" or when the caller supplies an ABC score.
  3. Semantic generation (AR) — generate_semantic(). The AR branch generates codec tokens from CODEC_OFFSET (151853) out of a 32,768-entry codebook, terminating on MUSIC_END (151852). Defaults temperature 1.0, top_p 0.95, top_k 100, repetition_penalty 1.2 over a 50-token window, min_tokens 200, max_tokens 9000. Semantic CFG guidance is 1.0 for full/melody (no negative branch) and 1.01 for off (a negative branch is built). Decode runs under CUDA graphs when available (cuda_graph.py), falling back to eager.
  4. Acoustic flow matching (NAR) — nar.py. song_chunks() draws the whole noise tensor on CPU in FP32 and slices it into chunks of size = (24576 - prefix_tokens - 3) // 2 frames. Each chunk gets one AR prefill (CachedNAR._prefill) to build a frozen KV cache, then a 32-step midpoint ODE (yue2_generation_config.json: ode_steps: 32, ode_method: "midpoint"). Midpoint means two velocity evaluations per step — 64 NAR forward passes per chunk. synthesize() returns CPU FP32 latents.
  5. Decoding — decode(). The MoT is moved to CPU, the VAE is loaded onto the GPU, and decode_tiled() runs in chunks of vae_core_frames with a 16-frame halo, returning audio to CPU. The VAE is then moved back to CPU.

Two structural facts drive everything below. First, 64 NAR forward passes per chunk dominate the FLOP budget — the AR stages are comparatively cheap and memory-bound. Second, the runtime is built around staging weights in and out of VRAM: decode() evicts the 6.76 GiB backbone to host RAM before the VAE runs. That design assumes the backbone fits in the first place.

Memory analysis — the disqualifying constraint

pipeline.py:158-165 is the whole story:

if self.device.type == "cuda":
    if not torch.cuda.is_bf16_supported():
        raise RuntimeError("The unquantized preset requires CUDA BF16 support")
    total = torch.cuda.get_device_properties(self.device).total_memory
    budget = min((self.memory_budget_gib - 2) * 2**30, total - 2 * 2**30)
    torch.cuda.set_per_process_memory_fraction(min(budget / total, 1), self.device)

The allocator is capped at min(memory_budget - 2 GiB, total - 2 GiB). Because the default memory_budget_gib is 24, the second term always wins on a small card: the usable ceiling is VRAM - 2 GiB, full stop. Raising memory_budget_gib cannot lift it.

GPU VRAM usable cap (total − 2 GiB) BF16 weights outcome
RTX 3050 4 GB Laptop 4 GiB 2.00 GiB 6.76 GiB ❌ OOM at load, short by 4.76 GiB
RTX 3050 6 GB Laptop (Victoria) 6 GiB 4.00 GiB 6.76 GiB ❌ OOM at load, short by 2.76 GiB
RTX 3050 6 GB desktop 6 GiB 4.00 GiB 6.76 GiB ❌ OOM at load, short by 2.76 GiB
RTX 3050 8 GB desktop 8 GiB 6.00 GiB 6.76 GiB ❌ OOM at load, short by 0.76 GiB

Even discounting the cap, the checkpoint would not fit. Weights alone are 6.76 GiB against 6 GiB of physical VRAM on Victoria, before the KV cache, before activations, before the ~0.3–0.5 GiB CUDA context and cuBLAS workspace.

The KV cache is statically sized in sampling.py:76-79 as max_seq_len = len(prefix) + sampling.max_tokens, at 28 layers × 8 KV heads × 128 head_dim × 2 (K and V) × 2 bytes = 112 KiB per token:

sequence tokens KV cache, BF16
semantic phase, ABC-conditioned prefix + 9000 cap 13,150 1.40 GiB
a 3.6-min song, NAR chunk (ar_length + nar_length) 14,895 1.55 GiB
full 24,576 context 24,576 2.62 GiB

The model card's own measurements agree with this reading. It reports 11.18 GiB peak VRAM for a typical 3.6-minute song on a 24 GB RTX 4090, and 14.08 GiB in maximum-context testing — roughly 4.4 GiB above the weights, consistent with the KV cache plus activations plus workspace. That is why the stated requirement is "24GB NVIDIA GPU with BF16 support" even though the weights are 6.76 GiB: the runtime wants the whole working set resident. Victoria has 6 GiB. The gap is not closable by tuning.

Compute and latency analysis

The scaling model is anchored on the one published measurement that matters: 71.04 s of wall clock for 214.85 s of audio on an RTX 4090, at a reported 139.48 LM tokens/s and 11.18 GiB peak. 71.04 × 139.48 ≈ 9,909 output tokens, which matches an ABC score plus ~5,371 semantic tokens for that song (5,371 frames × 40 ms = 214.8 s ✓). That gives a token count to work with.

Stage 1 — AR decode is memory-bound. Per token the AR path must stream the AR branch weights (1,409.4 M) plus the untied LM head (378.3 M), = 3.575 GB at BF16. On the 4090's 1008 GB/s that is 3.55 ms/token, so 9,909 tokens ≈ 35.1 s — about half the published 71.04 s, leaving ~35.9 s for the NAR and VAE.

Stage 2 — the NAR solver is compute-bound. For the same song the chunker produces one chunk (size = 10,211 ≥ 5,371 frames), with nar_length = 5,373 and ar_length = 9,522. Per velocity evaluation: 2 × 5,373 × 1,409.4 M = 15.14 TFLOP of linear work plus 0.66 TFLOP of attention. Times 64 evaluations = ≈1,011 TFLOP.

That budget pins the 4090's implied NAR rate at ≥28 TFLOP/s sustained over the residual 35.9 s — roughly 17% of the 4090's BF16 tensor peak, which is a realistic efficiency for a 2,048-wide model at this batch size. The model is therefore calibrated, not guessed.

Stage 3 — extrapolating to Victoria. Scaling the two stages independently, by bandwidth for the AR stage and by achievable tensor throughput for the NAR stage:

stage RTX 4090 (measured anchor) RTX 3050 6 GB Laptop (projected)
AR decode, 9,909 tokens 35.1 s @ 1008 GB/s 246.0 s @ 144 GB/s
NAR, 1,011 TFLOP ~35.9 s (≥28 TFLOPS implied, ≈17% MFU) 266 s – 674 s @ 1.5–3.8 TFLOPS (10–25% MFU)
decoded audio 214.85 s 214.85 s (does not fit — hypothetical)
end-to-end 71.04 s (RTF 0.33) 512 s – 920 s (RTF 2.38 – 4.28)

The NAR band applies the 4090's own implied 17% efficiency to Victoria's ~15.3 TFLOPS BF16 peak, widened to 10–25% because a 20-SM part with a 2 MB L2 is worse at keeping tensor cores fed than a 128-SM part with 72 MB. The central estimate is ~2.6 TFLOPS, giving ~389 s of NAR and ~635 s end-to-end.

The AR stage is a hard floor: 3.575 GB of weights per token cannot be cached on a 6 GB card, and 144 GB/s is a bandwidth fact. Even if the checkpoint fit, Victoria would run at 2.4–4.3× slower than real time — roughly 8.5–15 minutes for a 3.6-minute song, against 71 seconds on the 4090. The NAR band is the narrower of the two uncertainties now that the BF16 rate is settled (below); the AR floor is arithmetic.

An important secondary observation: for songs shorter than the 24,576-token context the NAR is not attention-bound. At nar_length 5,373 the attention term is 0.66 TFLOP against 15.14 TFLOP of linear work — 4%. Attention only reaches parity at sequence lengths around ten times the model's maximum. Optimising attention kernels for YuE2 on consumer hardware would be misdirected; the linear layers and the 64-step solver are the target.

BF16 on GA10x — checked, and it is not the problem

pipeline.py:159 gates on torch.cuda.is_bf16_supported(), which returns True for any compute capability ≥ 8.0, so sm_86 passes. That check is a capability probe rather than a throughput probe, and the surrounding folklore says GA10x runs BF16 at ~1/64 rate. That folklore is wrong for sm_86, and it was worth checking because it governs the whole NAR estimate.

For Victoria that works out to 2560 cores × 2 × 1.49 GHz ≈ 7.6 TFLOPS FP32, and a dense BF16 tensor peak of roughly 15.3 TFLOPS (2× FP32, FP32-accumulate). BF16 is a legitimate dtype here, not a cliff. The useful corollary is that BF16 and FP16 are the same speed on this GPU — so unlike Stable Audio 3, where the autoencoder's bf16-versus-fp16 choice is a real concern, YuE2's mandatory BF16 costs nothing relative to FP16.

YuE2 also offers no escape to FP16 even if it were faster: the checkpoint ships BF16, pipeline.py:217 loads it with torch_dtype=torch.bfloat16, and quantization.py:76-77 refuses to quantise anything that is not already the original BF16. The dtype is not negotiable — it just happens not to matter.

The compression escape routes, and why they do not rescue Victoria

FP8 AR quantisation — rejected by the hardware gate. quantization.py is the only built-in shrink, and its module docstring is explicit: "FP8 kernels require NVIDIA compute capability 8.9 or newer." The runtime enforces it (quantization.py:71-72):

if torch.cuda.get_device_capability(device) < (8, 9):
    raise RuntimeError("Experimental FP8 AR requires CUDA compute capability >=8.9")

GA107 is sm_86. The path is unavailable, and even where available the module self-describes as "quality_validation": "unvalidated", "performance_validation": "unvalidated", quantises only the AR linears, and restores exact BF16 before the NAR — so it would not have addressed the NAR cost anyway.

GGUF via audio.cpp — plausible but unvalidated. audio-cpp/Yue2-3B-GGUF publishes Q8_0 (4,264 MB) and Q4_0 (2,666 MB) main models plus F16 (265 MB) and F32 (530 MB) VAEs, converted from upstream revision 1a96eca6. Q4_0 + F16 VAE is 2.94 GB of weights, which does fit 6 GiB. But:

So the honest position is: the reference runtime is out of reach, and the only path that could fit is an unvalidated third-party quantisation that would still be several times slower than real time.

Verdict

question answer
Does YuE2-3B run on Victoria with the released runtime? No. 6.76 GiB of BF16 weights against a hard 4.00 GiB allocator cap.
Does it run on any RTX 3050? No. The 8 GB desktop card caps at 6.00 GiB, still short of 6.76 GiB.
Can memory settings be tuned to fit? No. The cap is total − 2 GiB; memory_budget_gib cannot raise it.
Can it be quantised to fit? Not with the released code. FP8 requires sm_89+; GA107 is sm_86.
If it somehow fit, would it be usable? Marginally. 2.4–4.3× slower than real time — 8.5–15 min per 3.6-min song.
Is there any viable path at all? Only GGUF/audio.cpp Q4_0, unofficial, ~2.94 GB of weights, ~9 min per song, unvalidated.
What hardware would actually run it? A 24 GB card — RTX 3090/4090. The RTX PRO 6000 (96 GB Blackwell) already in the benchmark pool runs it with large headroom, and being sm_120 it also unlocks the FP8 path.

Recommendation: do not attempt YuE2 on Victoria. The 96 GB RTX PRO 6000 already used for the Whistle benchmarks is the correct host — 96 GiB against a 14.08 GiB worst-case working set is comfortable, and sm_120 clears the FP8 gate that sm_86 fails. If local 6 GB generation is genuinely required, the only candidate is the GGUF route at ~9 min/song; the reference PyTorch pipeline needs 24 GB.

What to measure next

Victoria was unreachable, so every number above is analytic. When it is back, these are the commands that replace estimates with measurements. The first two settle the memory question in under a minute and cost nothing.

ssh victoria

# 1. Confirm the SKU and the real ceiling. Expect 6144 MiB and a 4.00 GiB allocator cap.
nvidia-smi --query-gpu=name,memory.total,memory.free,driver_version,compute_cap --format=csv
nvidia-smi -q -d MEMORY | head -20

# 2. Confirm the two dtype facts the whole analysis turns on, before downloading 7 GB.
python3 - <<'PY'
import torch
print("torch", torch.__version__, "cuda", torch.version.cuda)
print("is_bf16_supported() ->", torch.cuda.is_bf16_supported())   # expected True (capability check only)
print("capability          ->", torch.cuda.get_device_capability())  # expected (8, 6) -> FP8 unavailable
a = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
b = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
c = torch.randn(8192, 8192, device="cuda", dtype=torch.float16)
d = torch.randn(8192, 8192, device="cuda", dtype=torch.float16)
def bench(x, y, n=50):
    for _ in range(5): x @ y
    torch.cuda.synchronize(); s = torch.cuda.Event(True); e = torch.cuda.Event(True)
    s.record()
    for _ in range(n): x @ y
    e.record(); torch.cuda.synchronize()
    return (2 * 8192**3 * n) / (s.elapsed_time(e) / 1e3) / 1e12
bf, fp = bench(a, b), bench(c, d)
print(f"bf16 {bf:8.2f} TFLOPS   fp16 {fp:8.2f} TFLOPS   ratio {bf/fp:.3f}")
# This ratio is the single number that decides the NAR band in the table above.
PY

# 3. Only if the above is encouraging: prove the OOM rather than assuming it.
python3 -m pip install "huggingface-hub==0.36.2"
hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .
python3 -m pip install ./yue2_infer-0.1.5-py3-none-any.whl
python3 - <<'PY'
import torch
from yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda")   # expect OOM / allocator refusal here
print(torch.cuda.max_memory_allocated() / 2**30, "GiB")
PY

Beyond that, in priority order:

  1. Confirm the BF16:FP16 GEMM ratio from step 2. The documentary answer is settled — BF16 on sm_86 is native and runs at the fp16-with-FP32-accumulate rate, so the two should measure ~1.0 — but a cuBLAS build that routes bf16 elsewhere would invalidate the NAR band above, and the measurement is nearly free. Report the ratio alongside achieved FP32 TFLOPS so the 15.3 TFLOPS peak assumption is anchored too.
  2. If a larger card becomes available, measure the stage split properly. Wrap plan(), generate_semantic(), synthesize() and decode() in CUDA events and compare against torch.cuda.max_memory_allocated(). The published 71.04 s figure does not decompose; the 35.1 s / 35.9 s split above is inferred from a bandwidth argument and should be replaced with instrumentation.
  3. Test the solver-step trade directly. num_inference_steps is exposed by the GGUF runtime and, in the reference path, ode_steps is validated to equal the midpoint protocol. Any reduction from 32 steps is a quality decision, not just a speed one — compare audio, not just latency, before claiming a win.
  4. Quantify offload_ar before trusting it. It swaps branch weights through host RAM rather than keeping one branch resident. Measure the PCIe cost of the swap at 9,522-token prefill; it may be that a VRAM-resident AR branch plus a streamed NAR branch is the better split.
  5. Do not pursue attention kernels for YuE2 on this hardware. At 4% of NAR FLOPs at realistic song lengths there is nothing to win, and the Stable Audio 3 analysis reached the same conclusion from the opposite direction — FFN and projections dominate, not attention.