chatterbox — architecture, model map & inference hot paths
A complete systems inspection of Resemble AI's Chatterbox speech synthesis and voice conversion suite. This report maps every user-facing class and backing neural module, clarifies the fundamental hybrid AR + Diffusion paradigm, contrasts the newly released Chatterbox-Flash with the reference repository, traces execution hot paths down to tensor shapes and hardware passes, collects empirical benchmarks, and dissects the major latency sinks blocking sub-100ms real-time generation.
01Overview & The Core Paradigm: AR vs. Diffusion
A common point of confusion is whether modern systems like Chatterbox are "purely Autoregressive" or "purely Diffusion". The answer is that Chatterbox is a multi-stage hybrid: an Autoregressive (AR) language modeling stage that feeds into a Continuous Flow-Matching (Diffusion) stage, which in turn feeds into an iSTFT neural vocoder.
Stage 1: Discrete Autoregression (T3): Solves the semantic alignment and duration problem. Text length is arbitrary, and speech cadence varies dynamically. An autoregressive transformer (LLaMA or GPT-2) generates discrete acoustic tokens (from an 8,192 vocabulary, at 25 tokens per second) strictly one token at a time, left-to-right, terminating on a [STOP] token.
Stage 2: Continuous Flow Matching (S3Gen): Solves the acoustic texture and timbre problem. Generating high-resolution continuous waveforms or 80-band spectrograms directly with an AR transformer causes context explosion. Instead, Stage 2 takes the 25 Hz discrete tokens, upsamples them to 50 Hz, and uses an Ordinary Differential Equation (Euler ODE) to transform random Gaussian noise into an 80-band mel-spectrogram over multiple integration steps (10 steps in Base, 2 steps in Turbo).
Stage 3: Deterministic Neural Vocoding (HiFT-GAN): Takes the 80-band mel-spectrogram, predicts pitch (F0), synthesizes harmonic sine waves, upsamples by 120× using transposed convolutions, and applies inverse STFT (4× hop) to synthesize 24,000 Hz waveform samples.
Normalized text + prompt audio (16 kHz + 24 kHz)
Autoregressive Transformer predicting 25 Hz speech tokens (1 token/step)
Conditional Flow Matching ODE solver (25 Hz tokens → 50 Hz mels)
Harmonic NSF + UpConv + iSTFT (mel → 24 kHz waveform) + Perth watermark
Across the official family, there are six distinct model flavors built around variations of this architecture:
| Model Flavor | Parameters | Target Domain | Stage 1 (Text-to-Token) | Stage 2 (Token-to-Mel) | Key Distinction |
|---|---|---|---|---|---|
| Chatterbox Base | ~500M | English zero-shot TTS | Autoregressive Llama 30L (Batch=2 CFG) | 10-step Euler CFM (20 U-Net passes w/ CFG) | Baseline reference quality, Perceiver prompt resampler |
| Chatterbox Multilingual (V2/V3) | ~500M | 23+ Languages zero-shot TTS | Autoregressive Llama 30L (Batch=2 CFG) | 10-step Euler CFM (20 U-Net passes w/ CFG) | Multi-script frontends (Cangjie5, Jamo, Hiragana), tail-trim |
| Chatterbox-Turbo | ~350M | Low-latency English agents | Autoregressive GPT-2 Med 24L (Batch=1, no CFG) | 2-step Euler MeanFlow (2 U-Net passes, no CFG) | Paralinguistic tags ([cough], [laugh]), 10× fewer flow steps |
| Chatterbox-Flash | ~500M | High-throughput streaming TTS | Block-Diffusion Masked Decoder (Llama 30L + [MASK]) |
Standard S3Gen Flow Matching | Generates 16–32 tokens in parallel per block; 9×–13× realtime |
| Chatterbox-Nano | ~110M | CPU / On-device edge TTS | Autoregressive GPT-2 Small 12L (Batch=1, no CFG) | 2-step Euler MeanFlow (2 U-Net passes, no CFG) | Lightweight edge target: 3× faster than realtime on 8 CPU cores |
| Chatterbox-VC | ~120M | Zero-shot Voice Conversion | None (Bypasses Stage 1 completely) | 10-step Euler CFM (or MeanFlow) | Source audio 16 kHz → S3Tokenizer → S3Gen vocoder directly |
02Codebase & Class Map
The codebase is organized under refs/chatterbox/src/chatterbox. The implementation cleanly separates the top-level
orchestration engines from the underlying modular neural networks:
Top-Level Orchestrators
ChatterboxTTS (in tts.py)ChatterboxTurboTTS (in tts_turbo.py)ChatterboxMultilingualTTS (in mtl_tts.py)ChatterboxVC (in vc.py)punc_norm), conditioning cache, and high-level generation.T3 Model Subsystem
T3 (in models/t3/t3.py)T3CondEnc & T3Cond (in cond_enc.py)T3HuggingfaceBackend (in t3_hf_backend.py)LearnedPositionEmbeddings (in learned_pos_emb.py)Perceiver (in perceiver.py)S3Gen Generator Subsystem
S3Gen / S3Token2Wav (in s3gen.py)UpsampleConformerEncoder (in upsample_encoder.py)CausalConditionalCFM (in flow_matching.py)ConditionalDecoder (in decoder.py)HiFTGenerator (in hifigan.py)ConvRNNF0Predictor (in f0_predictor.py)Acoustic & Text Frontends
VoiceEncoder (in voice_encoder.py)CAMPPlus (in xvector.py)S3Tokenizer (in s3tokenizer.py)EnTokenizer & MTLTokenizer (in tokenizers/)PerthImplicitWatermarker03Conditioning & Voice Cloning Pipeline
Zero-shot voice cloning in Chatterbox does not rely on a single global vector. Instead, it computes an ensemble of acoustic representations across two sample rates (16,000 Hz and 24,000 Hz):
| Extracted Feature | Input Rate | Extracting Module | Output Representation | Consuming Module |
|---|---|---|---|---|
| Speaker d-vector | 16 kHz | VoiceEncoder (3-layer LSTM, 768 units) |
Tensor (1, 256), L2-normalized |
T3CondEnc.spkr_enc → linear projection to T3 hidden size (1024) |
| Acoustic Prompt Tokens | 16 kHz | S3Tokenizer (25 tokens/sec) |
Tensor (1, 150) or (1, 375) in Turbo |
T3.speech_emb → embedded, then cross-attended via Perceiver (or direct in Turbo) |
| Speaker x-vector | 16 kHz | CAMPPlus (CAM++ backbone) |
Tensor (1, 192) → affine mapped to (1, 80) |
S3Gen.flow → injected into Conditional CFM estimator |
| Reference Mel Spectrogram | 24 kHz | S3Gen.mel_extractor (80 mel bands) |
Tensor (1, T_mel, 80) (up to 10s = 500 frames) |
S3Gen.flow → prepended as context condition for the diffusion ODE |
| Emotion Exaggeration | Scalar | Floating point argument exaggeration |
Tensor (1, 1, 1) |
T3CondEnc.emotion_adv_fc → linear projection to 1024-dim |
librosa.resample, running the LSTM VoiceEncoder,
tokenizing 6 seconds of speech with S3Tokenizer, and computing CAMPPlus embeddings takes ~150–350 ms.
All classes store this in self.conds = Conditionals(...). Once prepared, the conditionals can be saved to disk
(conds.pt) or reused across multiple generation requests without re-evaluating the prompt audio.
04Hot Path: Base & Multilingual (500M)
The baseline ChatterboxTTS.generate() and ChatterboxMultilingualTTS.generate() share an identical
underlying execution pipeline. The path traverses text normalization, T3 prefill, autoregressive token sampling with Classifier-Free Guidance,
conformer upsampling, 10-step Euler ODE flow matching, and iSTFTNet vocoding.
Step 1: Text Tokenization & Conditioning Prefill
# 1. Text normalization and Tokenization
text = punc_norm(text) # cleans quotes, dashes, appends period
text_tokens = tokenizer.text_to_tokens(text) # shape: (1, L_text)
# 2. Duplicate for Classifier-Free Guidance (CFG)
if cfg_weight > 0.0:
text_tokens = torch.cat([text_tokens, text_tokens], dim=0) # shape: (2, L_text)
# 3. Framing with BOT (255) and EOT (0) tokens
text_tokens = F.pad(text_tokens, (1, 0), value=sot)
text_tokens = F.pad(text_tokens, (0, 1), value=eot)
# 4. Prepare Embeddings
# cond_emb: (2, 1 + 32 + 1, 1024) [speaker, perceiver speech tokens, emotion]
# text_emb: (2, L_text + 2, 1024) [LearnedPositionEmbeddings added]
# For batch index 1 (unconditional branch): text_emb[1].zero_()
embeds = torch.cat([cond_emb, text_emb, bos_embed], dim=1) # shape: (2, L_total, 1024)
Step 2: T3 Autoregressive Decode Loop (Primary Bottleneck #1)
T3 uses a 30-layer Llama backbone (hidden_size=1024, intermediate_size=4096, num_heads=16, head_dim=64,
RoPE theta 500k). The prefill pass processes the prompt context and initializes the dynamic KV-cache. Then, the autoregressive
decode loop generates discrete speech tokens one by one (tokens 0..6560):
# Decode Loop in T3.inference():
past = output.past_key_values # Initialized from prefill
for i in range(max_new_tokens): # up to 1000 steps, typically 250-350 steps
# 1. Slice logits from last step: shape (2, 8194)
logits_step = output.logits[:, -1, :]
# 2. Classifier-Free Guidance combination (Batch 2 -> Batch 1)
cond = logits_step[0:1, :]
uncond = logits_step[1:2, :]
logits = cond + cfg_weight * (cond - uncond) # shape: (1, 8194)
# 3. Logits processing (Repetition Penalty, Min-P, Top-P)
logits = repetition_penalty_processor(generated_ids[:1, :], logits)
logits = logits / temperature
logits = min_p_warper(generated_ids[:1, :], logits)
logits = top_p_warper(generated_ids[:1, :], logits)
# 4. Sampling
probs = torch.softmax(logits, dim=-1)
next_token = torch.multinomial(probs, num_samples=1) # (1, 1) - CPU sync!
# 5. Check EOS token (6562)
if next_token.view(-1) == 6562: # CPU-GPU sync!
break
# 6. Embed next token with learned positional embedding
next_token_embed = self.speech_emb(next_token) + self.speech_pos_emb(i + 1)
next_token_embed = torch.cat([next_token_embed, next_token_embed], dim=0) # (2, 1, 1024)
# 7. Single-token Llama forward pass
output = self.patched_model(
inputs_embeds=next_token_embed,
past_key_values=past,
use_cache=True,
output_hidden_states=True # disables certain kernel fusions
)
past = output.past_key_values # dynamic tuple-of-tuples KV cache reallocation
Step 3: S3Gen Upsample & Conformer Encoding
The generated speech tokens (25 Hz) are filtered (dropping BOS/EOS/OOV tokens) and concatenated with the reference prompt speech tokens.
The sequence is passed through UpsampleConformerEncoder:
- Input:
(1, T_tokens)mapped to 512-dim embedding. - Conformer architecture: 6 blocks with
attention_heads=8,linear_units=2048, relative positional embeddings. - Upsampling: A 1D convolutional subsampling/transposed-conv layer upsamples temporal frames by 2× (from 25 tokens/sec to 50 mel frames/sec).
- Output projected to 80 channels:
mu = self.encoder_proj(h), shape:(1, 80, 2 * T_tokens).
Step 4: S3Gen Flow Matching (Primary Bottleneck #2)
In standard Chatterbox, the token-to-mel decoder uses CausalConditionalCFM.solve_euler() with n_timesteps=10.
Crucially, it enforces Classifier-Free Guidance at every diffusion step:
# Inside solve_euler() in flow_matching.py:
B, T = mu.size(0), x.size(2) # B=1, T = mel frames
# 10 Euler steps along cosine time schedule:
for t, r in zip(t_span[:-1], t_span[1:]):
# Form batch of size 2 (cond + uncond)
x_in[:B] = x_in[B:] = x
mu_in[:B] = mu # mu_in[B:] = 0 (uncond)
spks_in[:B] = spks # spks_in[B:] = 0 (uncond)
cond_in[:B] = cond # cond_in[B:] = 0 (uncond)
# Forward pass of ConditionalDecoder (U-Net Estimator) at BATCH SIZE 2:
dxdt = self.estimator(x=x_in, mask=mask_in, mu=mu_in, t=t_in, spks=spks_in, cond=cond_in)
# CFG combination
dxdt, cfg_dxdt = torch.split(dxdt, [B, B], dim=0)
dxdt = (1.0 + inference_cfg_rate) * dxdt - inference_cfg_rate * cfg_dxdt
# Euler integration step
dt = r - t
x = x + dt * dxdt
Because each step evaluates the 20-block U-Net estimator at batch_size=2 across 10 steps, S3Gen executes
20 full passes of the U-Net per generation call. This accounts for ~1.2 to 1.8 seconds of GPU execution time.
Step 5: HiFT-GAN Neural Source Filter Vocoding
The generated mel-spectrogram (1, 80, T_mel) is converted to a 24 kHz waveform using HiFTGenerator:
ConvRNNF0Predictor: 5 weight-normed 1D conv layers predict fundamental pitch frequency (F0) from mel frames.SourceModuleHnNSF: Generates harmonic sine wave excitation from upsampled F0 plus voiced/unvoiced noise.- Transposed Convolutional Upsampling: 3 stages with rates
[8, 5, 3](cumulative 120×) combined with multi-receptive field ResBlocks (kernel sizes 3, 7, 11). - iSTFT Synthesis: Post-conv predicts STFT magnitude and phase, which are transformed via
torch.istft(hop length 4). - Total upsampling ratio:
120 * 4 = 480×(from 50 Hz mel to 24,000 Hz audio).
05Hot Path: Turbo (350M), Nano (110M) & Flash (500M)
Chatterbox-Turbo, Nano, and Flash represent two complementary philosophies for escaping the latency trap of the base model:
- Turbo & Nano tackle Stage 2: they retain sequential autoregressive decoding in Stage 1, but distill the Stage 2 flow-matching diffusion steps from 10 down to 2 using MeanFlow.
- Flash tackles Stage 1: it replaces sequential 1-token-at-a-time autoregression with a parallel Block-Diffusion Masked Decoder, generating entire blocks of tokens simultaneously.
| Architectural Axis | Chatterbox Base (500M) | Chatterbox-Turbo (350M) | Chatterbox-Nano (110M) | Chatterbox-Flash (500M) |
|---|---|---|---|---|
| Stage 1 Paradigm | Sequential Autoregressive (AR) | Sequential Autoregressive (AR) | Sequential Autoregressive (AR) | Block-Diffusion (Masked Discrete Diffusion) |
| Stage 1 Backbone | Llama 30L (1024-dim, 16h) | GPT-2 Medium 24L (1024-dim, 16h) | GPT-2 Small 12L (768-dim, 12h) | Llama 30L + [MASK] token (1024-dim) |
| Stage 1 Batch Size | 2 (CFG active) | 1 (no CFG) | 1 (no CFG) | 2 (PMI-CFG) |
| Token Parallelism | Strictly 1 token/step | Strictly 1 token/step | Strictly 1 token/step | 16 to 32 tokens in parallel per block |
| Position Embeddings | Learned pos emb per step | Native GPT-2 position encodings | Native GPT-2 position encodings | RoPE + block positional offsets |
| Paralinguistic Tags | No | Yes (native BPE tokens) | Yes (native BPE tokens) | Compatible |
| Stage 2 Flow Solver | 10-step Euler CFM w/ CFG | 2-step Euler MeanFlow | 2-step Euler MeanFlow | Standard S3Gen Flow Matching |
| Stage 2 Batch Size | 2 (CFG active) | 1 (no CFG) | 1 (no CFG) | 2 (CFG active) |
| Stage 2 U-Net Passes | 20 passes | 2 passes (10× reduction) | 2 passes (10× reduction) | 20 passes |
The MeanFlow Distillation Advantage in Turbo & Nano
In tts_turbo.py, S3Gen is initialized with meanflow=True and invokes basic_euler() with n_cfm_timesteps=2.
Because MeanFlow models are distilled directly against CFG outputs during post-training, no guidance is evaluated at inference time:
# Inside basic_euler() in flow_matching.py:
for t, r in zip(t_span[:-1], t_span[1:]):
t, r = t[None], r[None]
# Single forward pass at batch_size=1!
dxdt = self.estimator(x, mask=mask, mu=mu, t=t, spks=spks, cond=cond, r=r)
dt = r - t
x = x + dt * dxdt
Evaluating the U-Net estimator only twice at batch 1 (instead of 20 times at batch 2) drops S3Gen execution time from ~1500 ms down to ~60–90 ms on GPU. This removes Stage 2 as the primary serving bottleneck, leaving the T3 decode loop as the dominant latency sink.
Chatterbox-Flash: Parallel Masked Block-Diffusion
Rather than generating tokens one-by-one over hundreds of sequential Python iterations, Chatterbox-Flash reformulates Stage 1 token generation as Masked Discrete Diffusion (MDLM) structured into causal blocks:
- Block-by-Block Causal Streaming: Tokens are grouped into blocks of size $D = 16$, $24$, or $32$ tokens. Blocks are processed sequentially left-to-right, maintaining KV-cache history so that audio packets can stream immediately as each block finishes.
- Intra-Block Parallel Unmasking: Inside each block, all $D$ token slots begin as
[MASK]. Over $K \le 10$ diffusion denoising steps, the model predicts logits and unmasks multiple confident tokens simultaneously. - Adaptive Early Exit: An adaptive confidence schedule cuts denoising steps by ~20% ($\alpha = 0.5$ or $0.75$) with negligible acoustic difference.
- Inference Engine Architecture: Implemented via FlashInfer paged KV-cache and CUDA Graph capture, achieving up to 13× real-time (RTF = 0.076) with a Time to First Packet (TTFP) of just 103 ms.
Remaining Hazards in Turbo's T3 Decode Loop
While Turbo's architecture is theoretically fast, the reference PyTorch implementation in t3.py:inference_turbo() contains
several preventable software overheads:
- Repeated Tensor Concatenation: At every generated token,
input_ids = torch.cat(generated_speech_tokens, dim=1)allocates a brand new 2D tensor on GPU simply to pass intoRepetitionPenaltyLogitsProcessor. - Individual Embedding Lookups:
current_speech_embed = self.speech_emb(current_speech_token)issues a separate small GPU kernel launch per token rather than fusing or caching. - Host-Device Synchronization:
if torch.all(next_speech_token == self.hp.stop_speech_token): breaktriggers a blocking PCIe synchronization on every token, stalling the GPU pipeline.
06Hot Path: Voice Conversion (Chatterbox-VC)
ChatterboxVC provides zero-shot voice conversion without requiring text transcription or autoregressive decoding.
It operates strictly in the acoustic token domain:
16 kHz input speech waveform
Log-mel (100 fps) → VQ codebook (25 Hz speech tokens)
Conformer upsample → CFM ODE → iSTFTNet vocoder
24 kHz converted speech with target speaker timbre
Because T3 is skipped entirely, Voice Conversion does not suffer from autoregressive token decoding latency. Its latency is entirely determined by S3Tokenizer frame processing, 10-step CFM flow matching, and HiFT vocoding.
07Empirical Benchmarks & Serving Metrics
Benchmarking across vendor disclosures and independent community profiling reveals clear demarcations between the model tiers:
| Scenario / Configuration | Hardware Platform | T3 Sampling Rate | Time to First Audio (TTFA / TTFU) | Total Wall Time (10s audio) | Real-Time Factor (RTF) |
|---|---|---|---|---|---|
| Chatterbox Base (Non-streaming) | NVIDIA RTX 3090 (24 GB) | 47.67 tokens/sec | 7.62 s | 7.62 s | 0.682 (1.47× RT) |
| Chatterbox Base (100-tok chunk stream) | NVIDIA RTX 3090 (24 GB) | 34.45 tokens/sec | 3.18 s | 8.15 s | 0.816 (1.23× RT) |
| Chatterbox-Turbo (Vendor Claim) | Modern NVIDIA GPU | ~150+ tokens/sec | ~75 ms (chunked) | ~1.60 s | ≤ 0.160 (6.0× RT) |
| Chatterbox-Flash (D=16, α=0.5 default) | Modern NVIDIA GPU (FlashInfer) | Parallel block-diffusion | 118 ms | ~1.07 s | 0.107 (9.3× RT) |
| Chatterbox-Flash (D=32, α=0.75 max speed) | Modern NVIDIA GPU (FlashInfer) | Parallel block-diffusion | 103 ms | ~0.76 s | 0.076 (13.2× RT) |
| Chatterbox-Flash (Apple Silicon M4) | Apple M4 via MLX (4-bit QAT) | MLX Block-diffusion | — | ~6.65 s | 0.665 (1.5× RT) |
| Chatterbox Production Service | Commercial API cluster | Optimized engine | < 200 ms | — | — |
| Chatterbox-Nano (Edge / CPU) | 8 x86-64 CPU Cores | ~75 tokens/sec | — | ~3.33 s | ≤ 0.333 (3.0× RT) |
Latency Breakdown by Pipeline Component (Baseline Base Model, 10s Audio)
| Pipeline Stage | Operation | Batch Size | Approx. Wall Time | % of Total Latency |
|---|---|---|---|---|
| T3 Autoregressive Decode | 250 token steps through 30 Llama layers | 2 (CFG) | ~5,200 ms | 68.3% |
| S3Gen Flow Matching | 10 Euler ODE steps through 20 U-Net blocks | 2 (CFG) | ~1,550 ms | 20.3% |
| HiFT-GAN Vocoder | F0 prediction + 3-stage conv transpose + iSTFT | 1 | ~450 ms | 5.9% |
| Conditioning Extraction | VoiceEncoder LSTM + S3Tokenizer + CAMPPlus (uncached) | 1 | ~280 ms | 3.7% |
| Perth Watermarking & I/O | Neural watermarking on final waveform | 1 | ~135 ms | 1.8% |
08Bottlenecks & Optimization Ladder
To optimize Chatterbox into a high-throughput, ultra-low latency inference engine (analogous to what was achieved in dante
for Qwen3-ASR or kaminari for Qwen3-VL), several critical bottlenecks must be tackled:
Primary Bottlenecks
- HuggingFace Transformers Dynamic Cache: T3 relies on uncompiled HuggingFace Llama/GPT2 models with dynamically expanding KV-cache tuples, preventing CUDA Graph capture and triggering massive memory reallocations.
- Classifier-Free Guidance 2× Penalty: Running batch size 2 for T3 and CFM doubles the memory bandwidth and compute requirements.
- Repeated CPU-GPU Synchronization: Sampling logits on GPU, downloading token IDs to CPU for EOS detection, and updating repetition history via python lists creates latency bubbles.
- Multi-step CFM Evaluation: Base and Multilingual models run 20 passes of a deep U-Net; even Turbo runs 2 passes synchronously after the entire sentence finishes.
- Lack of Block-Causal Streaming: Generating audio only after the complete sentence finishes forces Time to First Utterance (TTFU) to equal the full generation time (3–7 seconds).
The Optimization Strategy
- Static KV-Cache &
torch.compile: Replace HuggingFace wrappers with a native minimal Llama/GPT-2 decoder using a pre-allocated static circular KV-cache. Compile the single-token step and capture into CUDA Graphs. - CFG Guidance Distillation: Adopt Turbo's guidance distillation approach for the multilingual and base models to drop batch size from 2 down to 1 across both T3 and S3Gen.
- GPU-Native Sampling & Ring Buffer: Move repetition penalty into a circular tensor on GPU; evaluate stop tokens without blocking CPU synchronization.
- Chunked Streaming Pipeline: Emit speech tokens from T3 into S3Gen in chunks of 25–50 tokens (1–2 seconds of audio), running CFM and vocoding concurrently while T3 continues decoding.
- Int8 / FP8 Weight Quantization: Quantize T3 and S3Gen linear projections to cut memory bandwidth demands on low-VRAM GPUs.
09Source Trail & Code References
| Subsystem / Metric | Source Reference |
|---|---|
| Engine entrypoints & conditionals | refs/chatterbox/src/chatterbox/tts.py, tts_turbo.py, mtl_tts.py, vc.py |
| T3 architecture & configs | refs/chatterbox/src/chatterbox/models/t3/t3.py, llama_configs.py, modules/t3_config.py |
| Conditional Flow Matching & MeanFlow | refs/chatterbox/src/chatterbox/models/s3gen/flow_matching.py, flow.py, decoder.py |
| HiFT-GAN & iSTFTNet Vocoder | refs/chatterbox/src/chatterbox/models/s3gen/hifigan.py, f0_predictor.py |
| Acoustic frontends | refs/chatterbox/src/chatterbox/models/voice_encoder/voice_encoder.py, xvector.py, s3tokenizer.py |
| Multilingual tokenizers | refs/chatterbox/src/chatterbox/models/tokenizers/tokenizer.py |
| RTX 3090 benchmarks & latency profile | Community empirical measurements in GitHub Issue #127 & Issue #193 |
| Turbo & Nano official specifications | Resemble AI official model cards, Demopage, and Podonos evaluation report |