lab

chatterbox — architecture, model map & inference hot paths

A complete systems inspection of Resemble AI's Chatterbox speech synthesis and voice conversion suite. This report maps every user-facing class and backing neural module, clarifies the fundamental hybrid AR + Diffusion paradigm, contrasts the newly released Chatterbox-Flash with the reference repository, traces execution hot paths down to tensor shapes and hardware passes, collects empirical benchmarks, and dissects the major latency sinks blocking sub-100ms real-time generation.

01Overview & The Core Paradigm: AR vs. Diffusion

A common point of confusion is whether modern systems like Chatterbox are "purely Autoregressive" or "purely Diffusion". The answer is that Chatterbox is a multi-stage hybrid: an Autoregressive (AR) language modeling stage that feeds into a Continuous Flow-Matching (Diffusion) stage, which in turn feeds into an iSTFT neural vocoder.

The Architectural Division of Labor

Stage 1: Discrete Autoregression (T3): Solves the semantic alignment and duration problem. Text length is arbitrary, and speech cadence varies dynamically. An autoregressive transformer (LLaMA or GPT-2) generates discrete acoustic tokens (from an 8,192 vocabulary, at 25 tokens per second) strictly one token at a time, left-to-right, terminating on a [STOP] token.

Stage 2: Continuous Flow Matching (S3Gen): Solves the acoustic texture and timbre problem. Generating high-resolution continuous waveforms or 80-band spectrograms directly with an AR transformer causes context explosion. Instead, Stage 2 takes the 25 Hz discrete tokens, upsamples them to 50 Hz, and uses an Ordinary Differential Equation (Euler ODE) to transform random Gaussian noise into an 80-band mel-spectrogram over multiple integration steps (10 steps in Base, 2 steps in Turbo).

Stage 3: Deterministic Neural Vocoding (HiFT-GAN): Takes the 80-band mel-spectrogram, predicts pitch (F0), synthesizes harmonic sine waves, upsamples by 120× using transposed convolutions, and applies inverse STFT (4× hop) to synthesize 24,000 Hz waveform samples.

Text / Conditioning
Normalized text + prompt audio (16 kHz + 24 kHz)
→
Stage 1: T3 LLM (Discrete AR)
Autoregressive Transformer predicting 25 Hz speech tokens (1 token/step)
→
Stage 2: S3Gen (Continuous CFM)
Conditional Flow Matching ODE solver (25 Hz tokens → 50 Hz mels)
→
Stage 3: HiFT Vocoder (iSTFTNet)
Harmonic NSF + UpConv + iSTFT (mel → 24 kHz waveform) + Perth watermark

Across the official family, there are six distinct model flavors built around variations of this architecture:

Model Flavor Parameters Target Domain Stage 1 (Text-to-Token) Stage 2 (Token-to-Mel) Key Distinction
Chatterbox Base ~500M English zero-shot TTS Autoregressive Llama 30L (Batch=2 CFG) 10-step Euler CFM (20 U-Net passes w/ CFG) Baseline reference quality, Perceiver prompt resampler
Chatterbox Multilingual (V2/V3) ~500M 23+ Languages zero-shot TTS Autoregressive Llama 30L (Batch=2 CFG) 10-step Euler CFM (20 U-Net passes w/ CFG) Multi-script frontends (Cangjie5, Jamo, Hiragana), tail-trim
Chatterbox-Turbo ~350M Low-latency English agents Autoregressive GPT-2 Med 24L (Batch=1, no CFG) 2-step Euler MeanFlow (2 U-Net passes, no CFG) Paralinguistic tags ([cough], [laugh]), 10× fewer flow steps
Chatterbox-Flash ~500M High-throughput streaming TTS Block-Diffusion Masked Decoder (Llama 30L + [MASK]) Standard S3Gen Flow Matching Generates 16–32 tokens in parallel per block; 9×–13× realtime
Chatterbox-Nano ~110M CPU / On-device edge TTS Autoregressive GPT-2 Small 12L (Batch=1, no CFG) 2-step Euler MeanFlow (2 U-Net passes, no CFG) Lightweight edge target: 3× faster than realtime on 8 CPU cores
Chatterbox-VC ~120M Zero-shot Voice Conversion None (Bypasses Stage 1 completely) 10-step Euler CFM (or MeanFlow) Source audio 16 kHz → S3Tokenizer → S3Gen vocoder directly

02Codebase & Class Map

The codebase is organized under refs/chatterbox/src/chatterbox. The implementation cleanly separates the top-level orchestration engines from the underlying modular neural networks:

Top-Level Orchestrators

ChatterboxTTS (in tts.py)
ChatterboxTurboTTS (in tts_turbo.py)
ChatterboxMultilingualTTS (in mtl_tts.py)
ChatterboxVC (in vc.py)
Responsible for check-pointing, device placement, text sanitization (punc_norm), conditioning cache, and high-level generation.

T3 Model Subsystem

T3 (in models/t3/t3.py)
T3CondEnc & T3Cond (in cond_enc.py)
T3HuggingfaceBackend (in t3_hf_backend.py)
LearnedPositionEmbeddings (in learned_pos_emb.py)
Perceiver (in perceiver.py)
Autoregressive LLM backbone. Translates text tokens into discrete speech tokens (vocab size 6561 + 2 special).

S3Gen Generator Subsystem

S3Gen / S3Token2Wav (in s3gen.py)
UpsampleConformerEncoder (in upsample_encoder.py)
CausalConditionalCFM (in flow_matching.py)
ConditionalDecoder (in decoder.py)
HiFTGenerator (in hifigan.py)
ConvRNNF0Predictor (in f0_predictor.py)
Speech generation backend. Translates tokens into mel spectrograms via CFM/MeanFlow, then vocodes to 24 kHz waveform.

Acoustic & Text Frontends

VoiceEncoder (in voice_encoder.py)
CAMPPlus (in xvector.py)
S3Tokenizer (in s3tokenizer.py)
EnTokenizer & MTLTokenizer (in tokenizers/)
PerthImplicitWatermarker
Extracts 256-d d-vectors, 192-d x-vectors, 25 Hz acoustic prompt tokens, multi-language orthography conversions, and watermark embedding.

03Conditioning & Voice Cloning Pipeline

Zero-shot voice cloning in Chatterbox does not rely on a single global vector. Instead, it computes an ensemble of acoustic representations across two sample rates (16,000 Hz and 24,000 Hz):

Extracted Feature Input Rate Extracting Module Output Representation Consuming Module
Speaker d-vector 16 kHz VoiceEncoder (3-layer LSTM, 768 units) Tensor (1, 256), L2-normalized T3CondEnc.spkr_enc → linear projection to T3 hidden size (1024)
Acoustic Prompt Tokens 16 kHz S3Tokenizer (25 tokens/sec) Tensor (1, 150) or (1, 375) in Turbo T3.speech_emb → embedded, then cross-attended via Perceiver (or direct in Turbo)
Speaker x-vector 16 kHz CAMPPlus (CAM++ backbone) Tensor (1, 192) → affine mapped to (1, 80) S3Gen.flow → injected into Conditional CFM estimator
Reference Mel Spectrogram 24 kHz S3Gen.mel_extractor (80 mel bands) Tensor (1, T_mel, 80) (up to 10s = 500 frames) S3Gen.flow → prepended as context condition for the diffusion ODE
Emotion Exaggeration Scalar Floating point argument exaggeration Tensor (1, 1, 1) T3CondEnc.emotion_adv_fc → linear projection to 1024-dim
Conditioning Cache Optimization Conditioning extraction is expensive: reading disk, resampling with librosa.resample, running the LSTM VoiceEncoder, tokenizing 6 seconds of speech with S3Tokenizer, and computing CAMPPlus embeddings takes ~150–350 ms. All classes store this in self.conds = Conditionals(...). Once prepared, the conditionals can be saved to disk (conds.pt) or reused across multiple generation requests without re-evaluating the prompt audio.

04Hot Path: Base & Multilingual (500M)

The baseline ChatterboxTTS.generate() and ChatterboxMultilingualTTS.generate() share an identical underlying execution pipeline. The path traverses text normalization, T3 prefill, autoregressive token sampling with Classifier-Free Guidance, conformer upsampling, 10-step Euler ODE flow matching, and iSTFTNet vocoding.

Step 1: Text Tokenization & Conditioning Prefill

# 1. Text normalization and Tokenization
text = punc_norm(text) # cleans quotes, dashes, appends period
text_tokens = tokenizer.text_to_tokens(text) # shape: (1, L_text)

# 2. Duplicate for Classifier-Free Guidance (CFG)
if cfg_weight > 0.0:
    text_tokens = torch.cat([text_tokens, text_tokens], dim=0) # shape: (2, L_text)

# 3. Framing with BOT (255) and EOT (0) tokens
text_tokens = F.pad(text_tokens, (1, 0), value=sot)
text_tokens = F.pad(text_tokens, (0, 1), value=eot)

# 4. Prepare Embeddings
# cond_emb: (2, 1 + 32 + 1, 1024) [speaker, perceiver speech tokens, emotion]
# text_emb: (2, L_text + 2, 1024) [LearnedPositionEmbeddings added]
# For batch index 1 (unconditional branch): text_emb[1].zero_()
embeds = torch.cat([cond_emb, text_emb, bos_embed], dim=1) # shape: (2, L_total, 1024)

Step 2: T3 Autoregressive Decode Loop (Primary Bottleneck #1)

T3 uses a 30-layer Llama backbone (hidden_size=1024, intermediate_size=4096, num_heads=16, head_dim=64, RoPE theta 500k). The prefill pass processes the prompt context and initializes the dynamic KV-cache. Then, the autoregressive decode loop generates discrete speech tokens one by one (tokens 0..6560):

# Decode Loop in T3.inference():
past = output.past_key_values # Initialized from prefill

for i in range(max_new_tokens): # up to 1000 steps, typically 250-350 steps
    # 1. Slice logits from last step: shape (2, 8194)
    logits_step = output.logits[:, -1, :]
    
    # 2. Classifier-Free Guidance combination (Batch 2 -> Batch 1)
    cond   = logits_step[0:1, :]
    uncond = logits_step[1:2, :]
    logits = cond + cfg_weight * (cond - uncond) # shape: (1, 8194)
    
    # 3. Logits processing (Repetition Penalty, Min-P, Top-P)
    logits = repetition_penalty_processor(generated_ids[:1, :], logits)
    logits = logits / temperature
    logits = min_p_warper(generated_ids[:1, :], logits)
    logits = top_p_warper(generated_ids[:1, :], logits)
    
    # 4. Sampling
    probs = torch.softmax(logits, dim=-1)
    next_token = torch.multinomial(probs, num_samples=1) # (1, 1) - CPU sync!
    
    # 5. Check EOS token (6562)
    if next_token.view(-1) == 6562: # CPU-GPU sync!
        break
        
    # 6. Embed next token with learned positional embedding
    next_token_embed = self.speech_emb(next_token) + self.speech_pos_emb(i + 1)
    next_token_embed = torch.cat([next_token_embed, next_token_embed], dim=0) # (2, 1, 1024)
    
    # 7. Single-token Llama forward pass
    output = self.patched_model(
        inputs_embeds=next_token_embed,
        past_key_values=past,
        use_cache=True,
        output_hidden_states=True # disables certain kernel fusions
    )
    past = output.past_key_values # dynamic tuple-of-tuples KV cache reallocation

Step 3: S3Gen Upsample & Conformer Encoding

The generated speech tokens (25 Hz) are filtered (dropping BOS/EOS/OOV tokens) and concatenated with the reference prompt speech tokens. The sequence is passed through UpsampleConformerEncoder:

  • Input: (1, T_tokens) mapped to 512-dim embedding.
  • Conformer architecture: 6 blocks with attention_heads=8, linear_units=2048, relative positional embeddings.
  • Upsampling: A 1D convolutional subsampling/transposed-conv layer upsamples temporal frames by 2× (from 25 tokens/sec to 50 mel frames/sec).
  • Output projected to 80 channels: mu = self.encoder_proj(h), shape: (1, 80, 2 * T_tokens).

Step 4: S3Gen Flow Matching (Primary Bottleneck #2)

In standard Chatterbox, the token-to-mel decoder uses CausalConditionalCFM.solve_euler() with n_timesteps=10. Crucially, it enforces Classifier-Free Guidance at every diffusion step:

# Inside solve_euler() in flow_matching.py:
B, T = mu.size(0), x.size(2) # B=1, T = mel frames

# 10 Euler steps along cosine time schedule:
for t, r in zip(t_span[:-1], t_span[1:]):
    # Form batch of size 2 (cond + uncond)
    x_in[:B] = x_in[B:] = x
    mu_in[:B] = mu             # mu_in[B:] = 0 (uncond)
    spks_in[:B] = spks         # spks_in[B:] = 0 (uncond)
    cond_in[:B] = cond         # cond_in[B:] = 0 (uncond)
    
    # Forward pass of ConditionalDecoder (U-Net Estimator) at BATCH SIZE 2:
    dxdt = self.estimator(x=x_in, mask=mask_in, mu=mu_in, t=t_in, spks=spks_in, cond=cond_in)
    
    # CFG combination
    dxdt, cfg_dxdt = torch.split(dxdt, [B, B], dim=0)
    dxdt = (1.0 + inference_cfg_rate) * dxdt - inference_cfg_rate * cfg_dxdt
    
    # Euler integration step
    dt = r - t
    x = x + dt * dxdt

Because each step evaluates the 20-block U-Net estimator at batch_size=2 across 10 steps, S3Gen executes 20 full passes of the U-Net per generation call. This accounts for ~1.2 to 1.8 seconds of GPU execution time.

Step 5: HiFT-GAN Neural Source Filter Vocoding

The generated mel-spectrogram (1, 80, T_mel) is converted to a 24 kHz waveform using HiFTGenerator:

  1. ConvRNNF0Predictor: 5 weight-normed 1D conv layers predict fundamental pitch frequency (F0) from mel frames.
  2. SourceModuleHnNSF: Generates harmonic sine wave excitation from upsampled F0 plus voiced/unvoiced noise.
  3. Transposed Convolutional Upsampling: 3 stages with rates [8, 5, 3] (cumulative 120×) combined with multi-receptive field ResBlocks (kernel sizes 3, 7, 11).
  4. iSTFT Synthesis: Post-conv predicts STFT magnitude and phase, which are transformed via torch.istft (hop length 4).
  5. Total upsampling ratio: 120 * 4 = 480× (from 50 Hz mel to 24,000 Hz audio).

05Hot Path: Turbo (350M), Nano (110M) & Flash (500M)

Chatterbox-Turbo, Nano, and Flash represent two complementary philosophies for escaping the latency trap of the base model:

  • Turbo & Nano tackle Stage 2: they retain sequential autoregressive decoding in Stage 1, but distill the Stage 2 flow-matching diffusion steps from 10 down to 2 using MeanFlow.
  • Flash tackles Stage 1: it replaces sequential 1-token-at-a-time autoregression with a parallel Block-Diffusion Masked Decoder, generating entire blocks of tokens simultaneously.
Architectural Axis Chatterbox Base (500M) Chatterbox-Turbo (350M) Chatterbox-Nano (110M) Chatterbox-Flash (500M)
Stage 1 Paradigm Sequential Autoregressive (AR) Sequential Autoregressive (AR) Sequential Autoregressive (AR) Block-Diffusion (Masked Discrete Diffusion)
Stage 1 Backbone Llama 30L (1024-dim, 16h) GPT-2 Medium 24L (1024-dim, 16h) GPT-2 Small 12L (768-dim, 12h) Llama 30L + [MASK] token (1024-dim)
Stage 1 Batch Size 2 (CFG active) 1 (no CFG) 1 (no CFG) 2 (PMI-CFG)
Token Parallelism Strictly 1 token/step Strictly 1 token/step Strictly 1 token/step 16 to 32 tokens in parallel per block
Position Embeddings Learned pos emb per step Native GPT-2 position encodings Native GPT-2 position encodings RoPE + block positional offsets
Paralinguistic Tags No Yes (native BPE tokens) Yes (native BPE tokens) Compatible
Stage 2 Flow Solver 10-step Euler CFM w/ CFG 2-step Euler MeanFlow 2-step Euler MeanFlow Standard S3Gen Flow Matching
Stage 2 Batch Size 2 (CFG active) 1 (no CFG) 1 (no CFG) 2 (CFG active)
Stage 2 U-Net Passes 20 passes 2 passes (10× reduction) 2 passes (10× reduction) 20 passes

The MeanFlow Distillation Advantage in Turbo & Nano

In tts_turbo.py, S3Gen is initialized with meanflow=True and invokes basic_euler() with n_cfm_timesteps=2. Because MeanFlow models are distilled directly against CFG outputs during post-training, no guidance is evaluated at inference time:

# Inside basic_euler() in flow_matching.py:
for t, r in zip(t_span[:-1], t_span[1:]):
    t, r = t[None], r[None]
    # Single forward pass at batch_size=1!
    dxdt = self.estimator(x, mask=mask, mu=mu, t=t, spks=spks, cond=cond, r=r)
    dt = r - t
    x = x + dt * dxdt

Evaluating the U-Net estimator only twice at batch 1 (instead of 20 times at batch 2) drops S3Gen execution time from ~1500 ms down to ~60–90 ms on GPU. This removes Stage 2 as the primary serving bottleneck, leaving the T3 decode loop as the dominant latency sink.

Chatterbox-Flash: Parallel Masked Block-Diffusion

Rather than generating tokens one-by-one over hundreds of sequential Python iterations, Chatterbox-Flash reformulates Stage 1 token generation as Masked Discrete Diffusion (MDLM) structured into causal blocks:

  • Block-by-Block Causal Streaming: Tokens are grouped into blocks of size $D = 16$, $24$, or $32$ tokens. Blocks are processed sequentially left-to-right, maintaining KV-cache history so that audio packets can stream immediately as each block finishes.
  • Intra-Block Parallel Unmasking: Inside each block, all $D$ token slots begin as [MASK]. Over $K \le 10$ diffusion denoising steps, the model predicts logits and unmasks multiple confident tokens simultaneously.
  • Adaptive Early Exit: An adaptive confidence schedule cuts denoising steps by ~20% ($\alpha = 0.5$ or $0.75$) with negligible acoustic difference.
  • Inference Engine Architecture: Implemented via FlashInfer paged KV-cache and CUDA Graph capture, achieving up to 13× real-time (RTF = 0.076) with a Time to First Packet (TTFP) of just 103 ms.

Remaining Hazards in Turbo's T3 Decode Loop

While Turbo's architecture is theoretically fast, the reference PyTorch implementation in t3.py:inference_turbo() contains several preventable software overheads:

  • Repeated Tensor Concatenation: At every generated token, input_ids = torch.cat(generated_speech_tokens, dim=1) allocates a brand new 2D tensor on GPU simply to pass into RepetitionPenaltyLogitsProcessor.
  • Individual Embedding Lookups: current_speech_embed = self.speech_emb(current_speech_token) issues a separate small GPU kernel launch per token rather than fusing or caching.
  • Host-Device Synchronization: if torch.all(next_speech_token == self.hp.stop_speech_token): break triggers a blocking PCIe synchronization on every token, stalling the GPU pipeline.

06Hot Path: Voice Conversion (Chatterbox-VC)

ChatterboxVC provides zero-shot voice conversion without requiring text transcription or autoregressive decoding. It operates strictly in the acoustic token domain:

Source Audio
16 kHz input speech waveform
→
S3Tokenizer
Log-mel (100 fps) → VQ codebook (25 Hz speech tokens)
→
S3Gen (Flow + HiFT)
Conformer upsample → CFM ODE → iSTFTNet vocoder
→
Target Audio
24 kHz converted speech with target speaker timbre

Because T3 is skipped entirely, Voice Conversion does not suffer from autoregressive token decoding latency. Its latency is entirely determined by S3Tokenizer frame processing, 10-step CFM flow matching, and HiFT vocoding.

07Empirical Benchmarks & Serving Metrics

Benchmarking across vendor disclosures and independent community profiling reveals clear demarcations between the model tiers:

Scenario / Configuration Hardware Platform T3 Sampling Rate Time to First Audio (TTFA / TTFU) Total Wall Time (10s audio) Real-Time Factor (RTF)
Chatterbox Base (Non-streaming) NVIDIA RTX 3090 (24 GB) 47.67 tokens/sec 7.62 s 7.62 s 0.682 (1.47× RT)
Chatterbox Base (100-tok chunk stream) NVIDIA RTX 3090 (24 GB) 34.45 tokens/sec 3.18 s 8.15 s 0.816 (1.23× RT)
Chatterbox-Turbo (Vendor Claim) Modern NVIDIA GPU ~150+ tokens/sec ~75 ms (chunked) ~1.60 s ≤ 0.160 (6.0× RT)
Chatterbox-Flash (D=16, α=0.5 default) Modern NVIDIA GPU (FlashInfer) Parallel block-diffusion 118 ms ~1.07 s 0.107 (9.3× RT)
Chatterbox-Flash (D=32, α=0.75 max speed) Modern NVIDIA GPU (FlashInfer) Parallel block-diffusion 103 ms ~0.76 s 0.076 (13.2× RT)
Chatterbox-Flash (Apple Silicon M4) Apple M4 via MLX (4-bit QAT) MLX Block-diffusion — ~6.65 s 0.665 (1.5× RT)
Chatterbox Production Service Commercial API cluster Optimized engine < 200 ms — —
Chatterbox-Nano (Edge / CPU) 8 x86-64 CPU Cores ~75 tokens/sec — ~3.33 s ≤ 0.333 (3.0× RT)

Latency Breakdown by Pipeline Component (Baseline Base Model, 10s Audio)

Pipeline Stage Operation Batch Size Approx. Wall Time % of Total Latency
T3 Autoregressive Decode 250 token steps through 30 Llama layers 2 (CFG) ~5,200 ms 68.3%
S3Gen Flow Matching 10 Euler ODE steps through 20 U-Net blocks 2 (CFG) ~1,550 ms 20.3%
HiFT-GAN Vocoder F0 prediction + 3-stage conv transpose + iSTFT 1 ~450 ms 5.9%
Conditioning Extraction VoiceEncoder LSTM + S3Tokenizer + CAMPPlus (uncached) 1 ~280 ms 3.7%
Perth Watermarking & I/O Neural watermarking on final waveform 1 ~135 ms 1.8%

08Bottlenecks & Optimization Ladder

To optimize Chatterbox into a high-throughput, ultra-low latency inference engine (analogous to what was achieved in dante for Qwen3-ASR or kaminari for Qwen3-VL), several critical bottlenecks must be tackled:

Primary Bottlenecks

  • HuggingFace Transformers Dynamic Cache: T3 relies on uncompiled HuggingFace Llama/GPT2 models with dynamically expanding KV-cache tuples, preventing CUDA Graph capture and triggering massive memory reallocations.
  • Classifier-Free Guidance 2× Penalty: Running batch size 2 for T3 and CFM doubles the memory bandwidth and compute requirements.
  • Repeated CPU-GPU Synchronization: Sampling logits on GPU, downloading token IDs to CPU for EOS detection, and updating repetition history via python lists creates latency bubbles.
  • Multi-step CFM Evaluation: Base and Multilingual models run 20 passes of a deep U-Net; even Turbo runs 2 passes synchronously after the entire sentence finishes.
  • Lack of Block-Causal Streaming: Generating audio only after the complete sentence finishes forces Time to First Utterance (TTFU) to equal the full generation time (3–7 seconds).

The Optimization Strategy

  • Static KV-Cache & torch.compile: Replace HuggingFace wrappers with a native minimal Llama/GPT-2 decoder using a pre-allocated static circular KV-cache. Compile the single-token step and capture into CUDA Graphs.
  • CFG Guidance Distillation: Adopt Turbo's guidance distillation approach for the multilingual and base models to drop batch size from 2 down to 1 across both T3 and S3Gen.
  • GPU-Native Sampling & Ring Buffer: Move repetition penalty into a circular tensor on GPU; evaluate stop tokens without blocking CPU synchronization.
  • Chunked Streaming Pipeline: Emit speech tokens from T3 into S3Gen in chunks of 25–50 tokens (1–2 seconds of audio), running CFM and vocoding concurrently while T3 continues decoding.
  • Int8 / FP8 Weight Quantization: Quantize T3 and S3Gen linear projections to cut memory bandwidth demands on low-VRAM GPUs.

09Source Trail & Code References

Subsystem / MetricSource Reference
Engine entrypoints & conditionalsrefs/chatterbox/src/chatterbox/tts.py, tts_turbo.py, mtl_tts.py, vc.py
T3 architecture & configsrefs/chatterbox/src/chatterbox/models/t3/t3.py, llama_configs.py, modules/t3_config.py
Conditional Flow Matching & MeanFlowrefs/chatterbox/src/chatterbox/models/s3gen/flow_matching.py, flow.py, decoder.py
HiFT-GAN & iSTFTNet Vocoderrefs/chatterbox/src/chatterbox/models/s3gen/hifigan.py, f0_predictor.py
Acoustic frontendsrefs/chatterbox/src/chatterbox/models/voice_encoder/voice_encoder.py, xvector.py, s3tokenizer.py
Multilingual tokenizersrefs/chatterbox/src/chatterbox/models/tokenizers/tokenizer.py
RTX 3090 benchmarks & latency profileCommunity empirical measurements in GitHub Issue #127 & Issue #193
Turbo & Nano official specificationsResemble AI official model cards, Demopage, and Podonos evaluation report