YuE2-3B on Victoria: architecture, inference path and RTX 3050 feasibility
Analysed 2026-09-14 against the released m-a-p/YuE2-3B
checkpoint (model repo created 2026-09-09, revision
1a96eca688d6ae5d7f0feb88573fec89920fcd19) and the shipped
inference wheel yue2_infer-0.1.5-py3-none-any.whl. The
target is Victoria, an NVIDIA GeForce RTX 3050 6GB Laptop GPU (GA107,
sm_86, 6144 MiB VRAM, 2560 CUDA cores at 1.24–1.49 GHz,
96-bit bus at 12 Gbps = 144 GB/s memory bandwidth).
Victoria was offline for the whole analysis (Tailscale
last saw it 15 h earlier; ssh victoria timed out on the
first attempt and was not retried), so every latency figure in this
report is analytical — derived from the released weights, the released
runtime source, and the model card's published RTX 4090 / H800
measurements. Nothing here is a measured 3050 number. The section "What
to measure next" contains the script that replaces the estimates with
measurements once Victoria is reachable.
Bottom line: YuE2-3B cannot run on Victoria, and the blocker
is memory, not speed. The BF16 MoT backbone alone is 6.76 GiB
of weights. The pipeline refuses to allocate more than
total_memory - 2 GiB, which is 4.00 GiB on a 6 GB
card and 6.00 GiB even on an 8 GB card — so the checkpoint does
not fit on any RTX 3050 variant, at any setting the released runtime
supports. The one built-in compression path (FP8) requires compute
capability ≥ 8.9 and is rejected on sm_86. Every other
route to fitting requires a different runtime (GGUF/audio.cpp), and the
arithmetic for that route says a 3.6-minute song would still take
roughly 8.5–15 minutes to render. The 96 GB RTX PRO 6000 already in the
benchmark pool is the host that actually fits.
Reproducibility
No GPU was available, so the analysis is a static one over released artifacts. Three sources were read directly rather than inferred:
- Exact parameter inventory from the
model.safetensorsheader, fetched with HTTP range requests against the Hugging Face CDN (the header is the first 73,184 bytes; the 7.26 GB body was never downloaded). Full tensor names, shapes and dtypes were enumerated from that header. - The released runtime source, from
yue2_infer-0.1.5-py3-none-any.whl(66,117 bytes, pure Python, 3,796 lines). The files that matter here areyue2/pipeline.py(allocation policy),yue2/nar.py(the flow-matching solver),yue2/protocol.py(fixed generation constants),yue2/sampling.py(KV-cache sizing) andyue2/quantization.py(the FP8 gate). - The model card's own throughput table
(
m-a-p/YuE2-3B), used as the calibration anchor for the scaling model.
Model card: https://huggingface.co/m-a-p/YuE2-3B. Code: https://github.com/multimodal-art-projection/YuE2.
Architecture
YuE2-3B is a Mixture-of-Transformers (MoT): one
28-layer backbone in which every layer carries two parallel
branches — an autoregressive branch (self_attn,
mlp) and a non-autoregressive branch
(nar_self_attn, nar_mlp) — plus a separate
Oobleck VAE decoder. The AR branch writes the song as text and discrete
tokens; the NAR branch runs flow matching to turn those tokens into
continuous latents; the VAE turns latents into stereo audio.
| property | value | source |
|---|---|---|
| total parameters | 3,630,684,224 (3.631 B) | summed from the safetensors header |
| checkpoint size, BF16 | 7,261,368,448 B = 7.26 GB = 6.763 GiB | header / weights_manifest.json |
| hidden size | 2048 | config.json |
| layers | 28 | config.json |
| attention heads / KV heads | 16 / 8 (GQA 2:1), head_dim 128 | config.json |
| intermediate size | 6144 (SwiGLU) | config.json |
| vocabulary | 184,704 (Qwen tiktoken base + 32,768 codec tokens) | config.json, protocol.py |
| max positions | 24,576 | config.json |
| AR branch weights | 1,409.4 M | header (sums over layers.*.self_attn,
layers.*.mlp) |
| NAR branch weights | 1,409.4 M | header (sums over layers.*.nar_*) |
| embeddings | 378.3 M (embed_tokens) |
header |
| LM head | 378.3 M, untied | header (tie_word_embeddings: false) |
| latent position table | 50.3 M (24576 × 2048) | header |
timestep embedder + vae2llm/llm2vae |
4.9 M | header |
The two branches are exactly equal in size (1,409.4 M each), which is
worth knowing: the NAR half of the checkpoint is dead weight
during the AR phase and vice versa. No released runtime
exploits that yet; nar.py's optional
offload_ar flag is the only nod to it, and it swaps branch
weights through host RAM rather than keeping only one resident.
The layer code contains a combined branch
(modeling_yue2.py:256-295) that computes both
branches and merges them with
torch.where(mask_3d, ar_out, nar_out). That path is not
used at inference — sampling.py takes the AR-only
else branch when ar_mask is None, and
nar.py:161-167 drives the NAR modules directly, bypassing
DecoderLayer.forward entirely. So there is no 2× waste in
the released pipeline, but any reimplementation that calls the layer
naively will pay it.
The audio tokeniser
The VAE config fixes the latent geometry:
downsampling_ratio: 1920 at 48 kHz stereo, i.e. one
latent frame = 40.0 ms of audio, 25 frames per second.
| quantity | value |
|---|---|
| latent frame rate | 25 Hz (40 ms/frame) |
| latent channels | 64 |
| maximum latent frames | 24,576 → 983 s ≈ 16.4 min |
| decoder architecture | Oobleck,
c_mults=[1,2,4,8,16,32],
strides=[2,2,4,4,5,6], channels 64, SnakeBeta,
final_tanh: false |
| decoder weights, FP32 | 530 MB
(m-a-p/YuE2-Vae/model.safetensors) |
| tiled decode | decode_core_frames 1024,
halo 16 (512 when
memory_budget_gib <= 12) |
The VAE is derived from stable-audio-tools
(a6ae0cdf, MIT) and NVIDIA's SnakeBeta — the wheel ships
both licences. It always runs in FP32;
pipeline.py:371 pins
"vae_dtype": "float32".
Verified inference path
Four stages, exposed separately as plan() →
generate_semantic() → synthesize() →
decode(). One pipe(...) call chains all
four.
- Prompt assembly —
protocol.py:108-126.SongRequest.text()concatenates a mode instruction,[Tags], the style string,[Lyrics]and the lyrics.token_prefixes()prependsEOD(151643) and appendsABC_START/ABC_END/MUSIC_START(151847/151848/151851). - Symbolic planning (AR) —
plan(), phase"abc". The AR branch generates an ABC melody-and-chord score. Defaultstemperature 0.7,top_p 0.9,top_k 30,max_tokens 4096. Skipped entirely whencot="off"or when the caller supplies an ABC score. - Semantic generation (AR) —
generate_semantic(). The AR branch generates codec tokens fromCODEC_OFFSET(151853) out of a 32,768-entry codebook, terminating onMUSIC_END(151852). Defaultstemperature 1.0,top_p 0.95,top_k 100,repetition_penalty 1.2over a 50-token window,min_tokens 200,max_tokens 9000. Semantic CFG guidance is 1.0 forfull/melody(no negative branch) and 1.01 foroff(a negative branch is built). Decode runs under CUDA graphs when available (cuda_graph.py), falling back to eager. - Acoustic flow matching (NAR) —
nar.py.song_chunks()draws the whole noise tensor on CPU in FP32 and slices it into chunks ofsize = (24576 - prefix_tokens - 3) // 2frames. Each chunk gets one AR prefill (CachedNAR._prefill) to build a frozen KV cache, then a 32-step midpoint ODE (yue2_generation_config.json:ode_steps: 32,ode_method: "midpoint"). Midpoint means two velocity evaluations per step — 64 NAR forward passes per chunk.synthesize()returns CPU FP32 latents. - Decoding —
decode(). The MoT is moved to CPU, the VAE is loaded onto the GPU, anddecode_tiled()runs in chunks ofvae_core_frameswith a 16-frame halo, returning audio to CPU. The VAE is then moved back to CPU.
Two structural facts drive everything below. First, 64 NAR
forward passes per chunk dominate the FLOP budget — the AR
stages are comparatively cheap and memory-bound. Second, the
runtime is built around staging weights in and out of VRAM:
decode() evicts the 6.76 GiB backbone to host RAM before
the VAE runs. That design assumes the backbone fits in the
first place.
Memory analysis — the disqualifying constraint
pipeline.py:158-165 is the whole story:
if self.device.type == "cuda":
if not torch.cuda.is_bf16_supported():
raise RuntimeError("The unquantized preset requires CUDA BF16 support")
total = torch.cuda.get_device_properties(self.device).total_memory
budget = min((self.memory_budget_gib - 2) * 2**30, total - 2 * 2**30)
torch.cuda.set_per_process_memory_fraction(min(budget / total, 1), self.device)
The allocator is capped at
min(memory_budget - 2 GiB, total - 2 GiB). Because the
default memory_budget_gib is 24, the second term
always wins on a small card: the usable ceiling is
VRAM - 2 GiB, full stop. Raising
memory_budget_gib cannot lift it.
| GPU | VRAM | usable cap
(total − 2 GiB) |
BF16 weights | outcome |
|---|---|---|---|---|
| RTX 3050 4 GB Laptop | 4 GiB | 2.00 GiB | 6.76 GiB | ❌ OOM at load, short by 4.76 GiB |
| RTX 3050 6 GB Laptop (Victoria) | 6 GiB | 4.00 GiB | 6.76 GiB | ❌ OOM at load, short by 2.76 GiB |
| RTX 3050 6 GB desktop | 6 GiB | 4.00 GiB | 6.76 GiB | ❌ OOM at load, short by 2.76 GiB |
| RTX 3050 8 GB desktop | 8 GiB | 6.00 GiB | 6.76 GiB | ❌ OOM at load, short by 0.76 GiB |
Even discounting the cap, the checkpoint would not fit. Weights alone are 6.76 GiB against 6 GiB of physical VRAM on Victoria, before the KV cache, before activations, before the ~0.3–0.5 GiB CUDA context and cuBLAS workspace.
The KV cache is statically sized in sampling.py:76-79 as
max_seq_len = len(prefix) + sampling.max_tokens, at 28
layers × 8 KV heads × 128 head_dim × 2 (K and V) × 2 bytes = 112
KiB per token:
| sequence | tokens | KV cache, BF16 |
|---|---|---|
| semantic phase, ABC-conditioned prefix + 9000 cap | 13,150 | 1.40 GiB |
a 3.6-min song, NAR chunk (ar_length + nar_length) |
14,895 | 1.55 GiB |
| full 24,576 context | 24,576 | 2.62 GiB |
The model card's own measurements agree with this reading. It reports 11.18 GiB peak VRAM for a typical 3.6-minute song on a 24 GB RTX 4090, and 14.08 GiB in maximum-context testing — roughly 4.4 GiB above the weights, consistent with the KV cache plus activations plus workspace. That is why the stated requirement is "24GB NVIDIA GPU with BF16 support" even though the weights are 6.76 GiB: the runtime wants the whole working set resident. Victoria has 6 GiB. The gap is not closable by tuning.
Compute and latency analysis
The scaling model is anchored on the one published measurement that matters: 71.04 s of wall clock for 214.85 s of audio on an RTX 4090, at a reported 139.48 LM tokens/s and 11.18 GiB peak. 71.04 × 139.48 ≈ 9,909 output tokens, which matches an ABC score plus ~5,371 semantic tokens for that song (5,371 frames × 40 ms = 214.8 s ✓). That gives a token count to work with.
Stage 1 — AR decode is memory-bound. Per token the AR path must stream the AR branch weights (1,409.4 M) plus the untied LM head (378.3 M), = 3.575 GB at BF16. On the 4090's 1008 GB/s that is 3.55 ms/token, so 9,909 tokens ≈ 35.1 s — about half the published 71.04 s, leaving ~35.9 s for the NAR and VAE.
Stage 2 — the NAR solver is compute-bound. For the
same song the chunker produces one chunk
(size = 10,211 ≥ 5,371 frames), with
nar_length = 5,373 and ar_length = 9,522. Per
velocity evaluation: 2 × 5,373 × 1,409.4 M = 15.14
TFLOP of linear work plus 0.66 TFLOP of
attention. Times 64 evaluations = ≈1,011 TFLOP.
That budget pins the 4090's implied NAR rate at ≥28 TFLOP/s sustained over the residual 35.9 s — roughly 17% of the 4090's BF16 tensor peak, which is a realistic efficiency for a 2,048-wide model at this batch size. The model is therefore calibrated, not guessed.
Stage 3 — extrapolating to Victoria. Scaling the two stages independently, by bandwidth for the AR stage and by achievable tensor throughput for the NAR stage:
| stage | RTX 4090 (measured anchor) | RTX 3050 6 GB Laptop (projected) |
|---|---|---|
| AR decode, 9,909 tokens | 35.1 s @ 1008 GB/s | 246.0 s @ 144 GB/s |
| NAR, 1,011 TFLOP | ~35.9 s (≥28 TFLOPS implied, ≈17% MFU) | 266 s – 674 s @ 1.5–3.8 TFLOPS (10–25% MFU) |
| decoded audio | 214.85 s | 214.85 s (does not fit — hypothetical) |
| end-to-end | 71.04 s (RTF 0.33) | 512 s – 920 s (RTF 2.38 – 4.28) |
The NAR band applies the 4090's own implied 17% efficiency to Victoria's ~15.3 TFLOPS BF16 peak, widened to 10–25% because a 20-SM part with a 2 MB L2 is worse at keeping tensor cores fed than a 128-SM part with 72 MB. The central estimate is ~2.6 TFLOPS, giving ~389 s of NAR and ~635 s end-to-end.
The AR stage is a hard floor: 3.575 GB of weights per token cannot be cached on a 6 GB card, and 144 GB/s is a bandwidth fact. Even if the checkpoint fit, Victoria would run at 2.4–4.3× slower than real time — roughly 8.5–15 minutes for a 3.6-minute song, against 71 seconds on the 4090. The NAR band is the narrower of the two uncertainties now that the BF16 rate is settled (below); the AR floor is arithmetic.
An important secondary observation: for songs shorter than
the 24,576-token context the NAR is not attention-bound. At
nar_length 5,373 the attention term is 0.66 TFLOP against
15.14 TFLOP of linear work — 4%. Attention only reaches parity at
sequence lengths around ten times the model's maximum. Optimising
attention kernels for YuE2 on consumer hardware would be misdirected;
the linear layers and the 64-step solver are the target.
BF16 on GA10x — checked, and it is not the problem
pipeline.py:159 gates on
torch.cuda.is_bf16_supported(), which returns
True for any compute capability ≥ 8.0, so
sm_86 passes. That check is a capability probe
rather than a throughput probe, and the surrounding folklore says GA10x
runs BF16 at ~1/64 rate. That folklore is wrong for sm_86, and
it was worth checking because it governs the whole NAR
estimate.
- NVIDIA's GA102 whitepaper lists BF16 tensor throughput as identical to FP16-with-FP32-accumulate in its spec tables — 71/142 TFLOPS on the RTX 3090 (Table 9), 59.5/119 on the RTX 3080 (Table 2), 40.6/81.3 on the RTX 3070 (Table 10).
- The CUDA C++ Programming Guide lists
__nv_bfloat16tensor support natively for "compute capabilities 8.0, 8.6 and 8.7", and the PTX ISA gives dense.bf16mmathe same K-depths as.f16(.m16n8k8,.m16n8k16) — but with fp32 accumulators only. BF16 therefore cannot reach fp16's 2× FP16-accumulate path; it sits at the FP16-with-FP32-accumulate rate. - The "1/64" number traces to the CUDA arithmetic-throughput table for compute capability 6.1 (Pascal GP10x), where 16-bit SIMT arithmetic runs at 2 ops/clk/SM against 128 FP32. It has no bearing on Ampere.
For Victoria that works out to 2560 cores × 2 × 1.49 GHz ≈ 7.6 TFLOPS FP32, and a dense BF16 tensor peak of roughly 15.3 TFLOPS (2× FP32, FP32-accumulate). BF16 is a legitimate dtype here, not a cliff. The useful corollary is that BF16 and FP16 are the same speed on this GPU — so unlike Stable Audio 3, where the autoencoder's bf16-versus-fp16 choice is a real concern, YuE2's mandatory BF16 costs nothing relative to FP16.
YuE2 also offers no escape to FP16 even if it were
faster: the checkpoint ships BF16, pipeline.py:217 loads it
with torch_dtype=torch.bfloat16, and
quantization.py:76-77 refuses to quantise anything that is
not already the original BF16. The dtype is not negotiable — it just
happens not to matter.
The compression escape routes, and why they do not rescue Victoria
FP8 AR quantisation — rejected by the hardware gate.
quantization.py is the only built-in shrink, and its module
docstring is explicit: "FP8 kernels require NVIDIA compute
capability 8.9 or newer." The runtime enforces it
(quantization.py:71-72):
if torch.cuda.get_device_capability(device) < (8, 9):
raise RuntimeError("Experimental FP8 AR requires CUDA compute capability >=8.9")
GA107 is sm_86. The path is unavailable, and even where
available the module self-describes as
"quality_validation": "unvalidated", "performance_validation": "unvalidated",
quantises only the AR linears, and restores exact BF16 before the NAR —
so it would not have addressed the NAR cost anyway.
GGUF via audio.cpp — plausible but unvalidated. audio-cpp/Yue2-3B-GGUF
publishes Q8_0 (4,264 MB) and Q4_0 (2,666 MB) main models plus F16 (265
MB) and F32 (530 MB) VAEs, converted from upstream revision
1a96eca6. Q4_0 + F16 VAE is 2.94 GB of
weights, which does fit 6 GiB. But:
- Its own measurements are on an RTX 5090, where Q4_0 + F16 VAE still peaked at 7,755 MiB and ran a 221 s song in 44.15 s (RTF 0.1997). The 7,755 MiB peak against 2.94 GB of weights shows how much KV cache and workspace the runtime reserves; the margin on a 6 GB card is thin and song-length dependent.
- It is a dev branch, described by its authors as "ready for community testing, validation, and optimization".
- Its examples use
num_inference_steps=8rather than the reference 32 — a 4× cut in NAR work, at unstated quality cost. - Extrapolating the 5090's RTF 0.1997 by bandwidth (1792 → 144 GB/s ≈ 12.4×) gives RTF ≈ 2.5, i.e. ~9 minutes for a 3.6-minute song. That lands at the optimistic end of the 8.5–15 minute band derived above from first principles — a useful independent cross-check, since the two estimates share no arithmetic.
So the honest position is: the reference runtime is out of reach, and the only path that could fit is an unvalidated third-party quantisation that would still be several times slower than real time.
Verdict
| question | answer |
|---|---|
| Does YuE2-3B run on Victoria with the released runtime? | No. 6.76 GiB of BF16 weights against a hard 4.00 GiB allocator cap. |
| Does it run on any RTX 3050? | No. The 8 GB desktop card caps at 6.00 GiB, still short of 6.76 GiB. |
| Can memory settings be tuned to fit? | No. The cap is total − 2 GiB;
memory_budget_gib cannot raise it. |
| Can it be quantised to fit? | Not with the released code. FP8 requires
sm_89+; GA107 is sm_86. |
| If it somehow fit, would it be usable? | Marginally. 2.4–4.3× slower than real time — 8.5–15 min per 3.6-min song. |
| Is there any viable path at all? | Only GGUF/audio.cpp Q4_0, unofficial, ~2.94 GB of weights, ~9 min per song, unvalidated. |
| What hardware would actually run it? | A 24 GB card — RTX 3090/4090. The RTX PRO 6000 (96 GB
Blackwell) already in the benchmark pool runs it with large
headroom, and being sm_120 it also unlocks the FP8
path. |
Recommendation: do not attempt YuE2 on Victoria. The
96 GB RTX PRO 6000 already used for the Whistle benchmarks is the
correct host — 96 GiB against a 14.08 GiB worst-case working set is
comfortable, and sm_120 clears the FP8 gate that
sm_86 fails. If local 6 GB generation is genuinely
required, the only candidate is the GGUF route at ~9 min/song; the
reference PyTorch pipeline needs 24 GB.
What to measure next
Victoria was unreachable, so every number above is analytic. When it is back, these are the commands that replace estimates with measurements. The first two settle the memory question in under a minute and cost nothing.
ssh victoria
# 1. Confirm the SKU and the real ceiling. Expect 6144 MiB and a 4.00 GiB allocator cap.
nvidia-smi --query-gpu=name,memory.total,memory.free,driver_version,compute_cap --format=csv
nvidia-smi -q -d MEMORY | head -20
# 2. Confirm the two dtype facts the whole analysis turns on, before downloading 7 GB.
python3 - <<'PY'
import torch
print("torch", torch.__version__, "cuda", torch.version.cuda)
print("is_bf16_supported() ->", torch.cuda.is_bf16_supported()) # expected True (capability check only)
print("capability ->", torch.cuda.get_device_capability()) # expected (8, 6) -> FP8 unavailable
a = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
b = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
c = torch.randn(8192, 8192, device="cuda", dtype=torch.float16)
d = torch.randn(8192, 8192, device="cuda", dtype=torch.float16)
def bench(x, y, n=50):
for _ in range(5): x @ y
torch.cuda.synchronize(); s = torch.cuda.Event(True); e = torch.cuda.Event(True)
s.record()
for _ in range(n): x @ y
e.record(); torch.cuda.synchronize()
return (2 * 8192**3 * n) / (s.elapsed_time(e) / 1e3) / 1e12
bf, fp = bench(a, b), bench(c, d)
print(f"bf16 {bf:8.2f} TFLOPS fp16 {fp:8.2f} TFLOPS ratio {bf/fp:.3f}")
# This ratio is the single number that decides the NAR band in the table above.
PY
# 3. Only if the above is encouraging: prove the OOM rather than assuming it.
python3 -m pip install "huggingface-hub==0.36.2"
hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .
python3 -m pip install ./yue2_infer-0.1.5-py3-none-any.whl
python3 - <<'PY'
import torch
from yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda") # expect OOM / allocator refusal here
print(torch.cuda.max_memory_allocated() / 2**30, "GiB")
PY
Beyond that, in priority order:
- Confirm the BF16:FP16 GEMM ratio from step 2. The documentary answer is settled — BF16 on sm_86 is native and runs at the fp16-with-FP32-accumulate rate, so the two should measure ~1.0 — but a cuBLAS build that routes bf16 elsewhere would invalidate the NAR band above, and the measurement is nearly free. Report the ratio alongside achieved FP32 TFLOPS so the 15.3 TFLOPS peak assumption is anchored too.
- If a larger card becomes available, measure the stage split
properly. Wrap
plan(),generate_semantic(),synthesize()anddecode()in CUDA events and compare againsttorch.cuda.max_memory_allocated(). The published 71.04 s figure does not decompose; the 35.1 s / 35.9 s split above is inferred from a bandwidth argument and should be replaced with instrumentation. - Test the solver-step trade directly.
num_inference_stepsis exposed by the GGUF runtime and, in the reference path,ode_stepsis validated to equal the midpoint protocol. Any reduction from 32 steps is a quality decision, not just a speed one — compare audio, not just latency, before claiming a win. - Quantify
offload_arbefore trusting it. It swaps branch weights through host RAM rather than keeping one branch resident. Measure the PCIe cost of the swap at 9,522-token prefill; it may be that a VRAM-resident AR branch plus a streamed NAR branch is the better split. - Do not pursue attention kernels for YuE2 on this hardware. At 4% of NAR FLOPs at realistic song lengths there is nothing to win, and the Stable Audio 3 analysis reached the same conclusion from the opposite direction — FFN and projections dominate, not attention.