kaminari — a single-GPU Qwen3-VL inference engine, explained
A full walkthrough of kaminari (雷, "thunder"): a minimal, custom
inference engine for Qwen3-VL built for OCR/VQA on a 6 GB laptop GPU. What every layer does,
how the request flows, where the 3.49× single-request and 9.35× batch speedups come from,
how it's tested without a GPU — and what it would take to swap Qwen3-VL for GLM-OCR.
01TL;DR
kaminari is a from-scratch VLM inference engine: a native PyTorch reimplementation
of the Qwen3-VL text decoder wrapped around the HF vision tower, a preallocated static KV cache, a
torch.compile-shaped one-token decode loop, offline weight-only int4 quantization via
torchao, and a FastAPI serving layer with three schedulers (microbatch, continuous slots, SSE
streaming). It runs Qwen3-VL-2B at 100.6 tok/s decode —
3.49× the HF generate() baseline — in 2.3 GB of VRAM
on an RTX 3050 (6 GB, mobile), and 9.35× aggregate HF throughput at batch 4.
engine
native Qwen3-VL port · fused QKV/gate-up · GQA via SDPA · 3D M-RoPE · deepstack vision features
2,783 lines of production Python
decode
static KV buffers mutated in place · fixed-rank positions · compiled one-token step
100.6 tok/s · 5.18 s / 512-token gen
memory
torchao weight-only int4 on decoder linears · offline sidecar pack · vision + lm_head stay bf16
2.3 GB peak (−45% vs bf16)
serving
FastAPI · microbatch / continuous-slot / streaming schedulers · SSE token stream
batch=2 sweet spot on 6 GB
The whole test suite (27 tests) runs on CPU against a ~1 MB synthetic checkpoint — no GPU needed to verify engine correctness. GPU-only paths are exercised by benchmark scripts on the real box.
02Why kaminari exists
The author runs OCR and local multimodal workloads (receipt transcription, page OCR, captioning,
VQA) on a mobile RTX 3050 with 6 GB of VRAM. The stock path — HuggingFace
Qwen3VLForConditionalGeneration.generate() — manages 28.8 tok/s on that card, spends
memory on dynamic cache growth, and provides no control over batching or preprocessing.
kaminari's design constraints follow from that hardware reality:
- every MB counts — the model plus a 1024-token KV cache must leave headroom for the 6 GB card's other tenants; int4 weight-only quantization plus a preallocated cache cut peak VRAM from 4.2 GB to 2.3 GB
- compile shapes must be stable — variable image sizes churn
torch.compileguards; fixed visual-token buckets pin the shapes - OCR doesn't need sampling — greedy-only decoding removes the sampler from the hot path entirely
- one GPU owner — a single-GPU runtime doesn't need vLLM's distributed machinery; it needs a tight, auditable loop
The design borrows deliberately from three references: kestrel (moondream's edge inference stack — fixed-shape decode, prefix thinking), gpt-fast (pytorch-native hot path, static KV, weight-only quantization), and nano-vllm (minimal scheduler/paged-cache architecture). kaminari is what falls out of applying those ideas to a multimodal model on a small card.
03Architecture overview
src/kaminari/
├── __init__.py # exports Engine
├── api.py # FastAPI server, click CLI, pydantic schemas, SSE
├── vram_utils.py # VRAM stat helpers
├── engine/
│ ├── types.py # EngineConfig, Batch, Live, DecodeSlot(+Pool), StreamEvent
│ ├── frontend.py # cpu preprocessing: bucket → chat template → processor → Batch
│ ├── processor.py # HF processor loading, dtype + visual-bucket resolution
│ ├── runtime.py # Engine: generate / generate_stream / prefill / decode / finalize
│ ├── scheduler.py # MicrobatchScheduler, ContinuousBatchScheduler, StreamingScheduler
│ └── utils.py # metrics, compiled decode wrapper, sampling, token commit
└── model/
├── qwen3vl.py # Qwen3VLForConditionalGeneration (native) + load_qwen3vl
├── modules.py # TextModel, DecoderLayer, TextAttention (fused QKV), MLP, Mrope
├── cache.py # StaticKVLayer / StaticKVCache (preallocated, slot-addressable)
└── utils.py # rope/masks, checkpoint loading, projection fusion, int4 sidecar
Layer responsibilities
| layer | owns | never touches |
|---|---|---|
api.py | transport: base64/path media → PIL, request schemas, endpoints, server lifecycle | tensors beyond PIL images |
engine/scheduler.py | admission: bounded queue, coalescing window, slot pool, future scattering | model internals |
engine/runtime.py | execution: prefill → decode loop → finalize, live state, metrics | HTTP concerns |
engine/frontend.py | cpu preprocessing: bucketing, chat template, processor invocation | cuda |
model/* | tensors: forward passes, cache, checkpoint loading, fusion, quantization restore | request concepts |
State model
Three immutable dataclasses carry the entire generation lifecycle, updated via
dataclasses.replace() — state changes flow through return values, not mutation:
Batch— one preprocessed request bundle: request ids, chat texts, padded tensor map from the processor, per-row prompt lengths, images, EOS/pad idsLive— hot decode state: static KV cache reference, accumulatedgenerated_token_ids,next_token_idcolumn, 3D M-RoPEposition_ids, runtimecache_position, per-row finished flagsDecodeSlot— aLivepinned to a physical cache row for the continuous scheduler; slots live in a fixedDecodeSlotPoolsizedmax_batch_size
The one sanctioned exception to immutability: the KV cache buffers themselves mutate in place — that's the point of a static cache.
04Request lifecycle
blocking path (/generate)
/generate→
scheduler: bounded queue + wait_ms coalescing→
microbatch forms (≤ max_batch_size)→
Engine.generate() in a worker thread→
futures scattered back per row
Inside Engine.generate(), five stages run in order:
1 · preprocess (cpu)
Each image is letterboxed onto a fixed visual-token bucket canvas: the frontend
picks the aspect-closest divisor grid for the token budget (e.g. 512 tokens → 16×32 →
896×448 px at 28 px/token for a wide image, 448×896 for a tall one), pastes the
LANCZOS-resized image centered on a black canvas, then renders the chat template and calls the HF
processor with padding="max_length" / truncation=True bounded by
kv_cache_max_len − max_new_tokens. The result is one Batch whose padded
tensors are trimmed to the batch max prompt length before anything moves to CUDA — pad tokens
must never reach the cache.
2 · prefill (one padded batch forward)
Trimmed inputs move to device; the native get_rope_index computes 3D M-RoPE
positions and rope_deltas on CPU (integer extraction stays off the GPU path); the
static cache is built; a single forward runs vision encode + text prefill with
logits_to_keep sliced to each row's last non-pad prompt token. The first token is
sampled with greedy argmax immediately — TTFT is the prefill, there is no separate
"first decode step".
3 · decode (compiled one-token loop)
Each step calls a torch.compile(mode="reduce-overhead", fullgraph=True)-wrapped
closure with fixed-rank tensors only: token ids (batch, 1),
positions (3, batch, 1), cache_position (1,), and the full preallocated
KV buffers. Greedy selection argmaxes the final logits row outside the compiled callable and stays
on-device until commit. Per-row EOS flips a finished flag; finished rows are padded with pad ids
so the batch shape never changes mid-generation. The loop runs until every row is finished or
max_new_tokens is hit.
4 · finalize
commit_tokens detaches GPU tensors → CPU → python lists, trimming each row at its
first EOS; decode_text batch-decodes through the tokenizer. Metrics (preprocessing,
prefill, decode, forward, sampling, finalize, TTFT, tok/s, img/s, VRAM) are assembled into
EngineMetrics and attached to every row.
streaming path (/generate_stream)
The streaming endpoint bypasses batching entirely: a StreamingScheduler runs
Engine.generate_stream() — a synchronous generator sharing the same prefill, then
yielding one StreamEvent per decode step — in a worker thread, bridging events to
the event loop through an asyncio.Queue. The API layer formats each event as
data: {json}\n\n (SSE), checks request.is_disconnected() every cycle,
and emits a final done event carrying the full text and metrics.
05The native model port
kaminari does not call HuggingFace's forward(). The text decoder is a native
reimplementation in model/; the vision tower is reused from transformers. The split
is deliberate: the decoder is where every decode step is spent and where compile/quantization
control is needed; the vision tower runs once per request during prefill and HF's implementation
is already fused and correct.
| component | implementation | why |
|---|---|---|
| vision tower | HF Qwen2_5_VL-family vision encoder, called from the native wrapper | runs once per request; deepstack multi-level features and window attention are intricate and not the hot path |
| text decoder | native TextModel: 28 layers, GQA, QK-norm, M-RoPE, SwiGLU MLP | full control of shapes, cache writes, fusion, quantization targets |
| attention | fused qkv_proj + torch.nn.functional.scaled_dot_product_attention with native GQA | one GEMM instead of three per layer; SDPA's GQA path avoids explicit KV replication |
| MLP | fused gate_up_proj → SiLU → elementwise → down_proj | halves projection launches; compile-friendly layout |
| projection fusion | checkpoint remap at load time (fuse_checkpoint_projections): Q/K/V → QKV, gate/up → gate-up | the checkpoint stays upstream-compatible; fusion is a load-time transform, not a fork |
| positions | 3D M-RoPE (temporal/height/width), computed per request via get_rope_index, cached as rope_deltas for decode | matches Qwen3-VL's multimodal position semantics exactly |
| KV cache | StaticKVCache: preallocated (max_batch, heads, max_seq, head_dim) buffers per layer, written via cache_position, read as active-prefix slices or full buffers in decode | no per-token torch.cat; stable shapes for compile; slot-addressable rows for continuous batching |
(batch,1), the cache,
positions (3,batch,1), cache_position (1,), and a (usually
None) mask — and returns one logits row. Everything else (mask construction from the
static cache length, cache writes, position increment) happens inside the native forward with
no python .item() reductions and no dynamic shapes. That contract is
what makes reduce-overhead + fullgraph=True hold across every step of
every request.06The optimization ladder
Single-request decode throughput on the RTX 3050 (Qwen3-VL-2B, 512-token generations,
1024-token KV budget, greedy). Each row is a measurable milestone, kept in
benchmarks_3vl/ with the exact run JSON:
| optimization | tok/s | latency | peak vram | vs hf |
|---|---|---|---|---|
HF generate() baseline | 28.82 | 17.77 s | — | baseline |
| kaminari HF runtime (early wrapper) | 25.38 | 20.40 s | 4173 MB | −12.0% |
| native impl + GQA + fused QKV | 30.31 | 16.98 s | 4248 MB | +5.2% |
| custom int4 decoder | 46.15 | 11.29 s | 2275 MB | +60.2% |
| torchao int4 | 51.88 | 9.96 s | 2275 MB | +80.0% |
| int4 + torch.compile + static KV | 100.60 | 5.18 s | 2275 MB | +249.1% |
| batch infer, batch=4 (aggregate) | 269.40 | 7.91 s | — | +834.8% |
what each step actually bought
- native port (+5%) — the win wasn't raw speed; it was control. Owning the forward unlocked the fused projections, the static cache, and compile. (The early HF-wrapper stage was honestly slower than stock — the table keeps that negative result.)
- int4 (+60→80%) — decode on a 3050 is memory-bandwidth-bound; halving weight bytes nearly doubles decode. VRAM dropped 45% in the same move.
- compile + static KV (+94% on top of int4) — the big one. Fixed shapes +
reduce-overhead(CUDA-graph-backed) removed per-step kernel-launch and python overhead; static buffers removed cache-growth allocations. Together they took 5.18 s per 512-token generation. - batching (+835% aggregate) — one padded prefill and one decode step serve N rows; aggregate throughput scales until the card goes VRAM-bound.
the concurrency curve
Concurrent /generate traffic with microbatching enabled surfaces the 3050's limits
honestly. At 512 output tokens per request:
| batch | images/sec | tok/s | mean latency |
|---|---|---|---|
| 1 | 0.13 | 71.30 | 7.50 s |
| 2 | 0.24 | 131.21 | 8.42 s |
| 4 | 0.22 | 121.38 | 18.05 s |
batch=2 is the sweet spot: batch=4 doubles latency and loses throughput
because operations become VRAM-bound. The merged EngineConfig ships
max_batch_size=2 as its default for exactly this reason.
07int4 sidecar quantization
kaminari treats quantization as an offline packaging step, not a runtime
option: scripts/quantize_int4.py loads the bf16 model on a CUDA machine, converts
only the language-model attention/MLP linears to torchao weight-only int4, and writes a
sidecar pack — the non-quantized tensors (vision tower, embeddings, lm_head,
norms) stay in a pruned model.safetensors, the packed decoder weights go to
qwen3-vl-int4-pack.pt.
why this shape
- selective targets — only decoder linears are quantized; the vision encoder
and lm_head stay bf16, so visual fidelity and the final logits projection keep full precision.
The packer validates that no module outside
model.language_model.layers.*attention/MLP projections was converted - low-memory restore — the engine loads the pruned fp weights, moves to CUDA, then installs the int4 wrappers and reads the sidecar; torchao's tile-packed int4 requires the module to already live on the target device, a subtle ordering constraint the loader handles (and tests pin)
- explicit, not implicit —
quantization="int4"fails hard unlessint4_pack_pathis provided; there is no silent fallback path that quantizes at load time
scripts/ocr_eval.py exists for exactly this — transcribe a fixed image set at bf16
vs int4 and score normalized edit distance against reference texts — but the GPU runs haven't
happened. It's the one experiment that would close the loop on the quantization story.08The serving layer
| endpoint | method | behavior |
|---|---|---|
/health | GET | load status (starting/ok/error), config echo |
/generate | POST | one request; goes through the scheduler as a batch of 1 |
/generate_batch | POST | explicit row batch; each row is a scheduler job so microbatching still decides grouping |
/generate_stream | POST | SSE: one data: {json} per token + final done with metrics; disconnect-aware |
three schedulers, one engine
MicrobatchScheduler(default) — boundedasyncio.Queue(backpressure viamax_queue_size), a short coalescing window after the first job so concurrent requests batch together, oneEngine.generate()per batch in a worker thread, futures scattered back to rows. Per-job scheduler metrics (queue wait, microbatch wait, batch id/size) ride alongside engine metrics in every response.ContinuousBatchScheduler— admits requests into freeDecodeSlots and ticks active slots one token at a time, finalizing each row the moment it finishes instead of waiting for the whole batch. The slot pool owns a shared static cache with per-row physical rows;copy_row_frommigrates a finished row's KV state.StreamingScheduler— no batching, no slots: one stream at a time, token events bridged from the engine thread to the event loop.
The engine itself stays synchronous and batch-shaped — all async concerns live in the
scheduler layer, and the API layer never touches tensors. Server defaults match the measured
sweet spot: int4 on, max_batch_size=2, 50 ms coalescing window,
kv_cache_max_len=2048.
09Correctness engineering
An inference engine that can only be tested on a GPU is an engine you can't refactor.
kaminari's suite runs entirely on CPU: scripts/create_tiny_qwen3_vl_checkpoint.py
synthesizes a ~1 MB Qwen3-VL-compatible checkpoint (tiny-but-structurally-exact config:
2 layers, 4 heads, the real processor and chat template) into
tests/fixtures/tiny_qwen3_vl_cpu, and every test runs the real engine path against it.
| test file | what it pins |
|---|---|
test_tiny_qwen3_vl_fixture.py | end-to-end: frontend input building, native model forward, batch generation, prefill/decode Live state invariants (cache position, token accumulation, completion predicate) |
test_api_server.py | health / generate / concurrent microbatch serving against a dummy runtime — transport correctness without weights |
test_decode_slots.py | static KV semantics: 1D position expansion per row, slot-row writes, active-prefix reads, copy_row_from |
test_quantization.py | int4 sidecar contracts: pack filename, module targeting, cuda-before-pack ordering, weight filtering |
test_runtime_compile.py | the compiled decode wrapper: decode_token closure, reduce-overhead + fullgraph=True |
test_utils_perf.py | VRAM stat helpers degrade gracefully on CPU |
27 tests, ruff check clean across src/ tests/ scripts/, deps pinned to
CPU wheels via [tool.uv.sources] so uv sync + pytest works on
any laptop. Testing locally is the whole story — no CI/CD, no remote runners.
10Repo state & roadmap
where it stands
- main carries everything: native engine, int4 sidecar, three schedulers,
SSE streaming, green CPU suite, ruff-clean lint, MIT license, committed benchmark history
(
benchmarks_3vl/JSONs back the README table) - parked as tags —
archive/qwen3-mtp(qwen3.5 adapter + MTP speculative decoding, 15 commits) andarchive/decode-slots-v1(slot exploration, 12 commits); restore withgit checkout <tag> - documented —
docs/project.mdmaps every module, the request path, the scripts inventory, and the simplification proposal;logs.mdcarries the full edit history
next unlocks, in the order that makes sense
| unlock | what it is | why it's next |
|---|---|---|
| int4 quality numbers | run scripts/ocr_eval.py bf16-vs-int4 on the OCR bench set | closes the one open scientific question; needs one GPU session |
| streaming verification + bench | GPU-verify /generate_stream, add streaming numbers to the README table | the streaming path merged untested on real hardware |
| core collapse | the simplification proposal: strip serving from the core, target 1–1.5k lines | the runtime should own one Live per call, not a request map |
| paged-kv extraction | standalone block allocator + block table → continuous batching without slot copies | systems-depth follow-up; a second project, not core work |
| prefix caching | reuse the fixed BOS+image-token prefix across OCR calls to the same page set | kestrel-style; large prefill savings for repeated OCR prompts |
11Swap analysis: replacing Qwen3-VL with GLM-OCR
GLM-OCR (zai-org, Jan 2026) is a purpose-built document-OCR VLM — and a genuinely interesting swap candidate for kaminari, because kaminari's stated hot path (page/receipt OCR) is exactly what it was trained for. This section is a fact-first look at what the model is, how it compares, and what an engine port would actually involve.
11.1 · what it is
| property | GLM-OCR | Qwen3-VL-2B (kaminari today) |
|---|---|---|
| total params | 0.9B vendor claim (1.33B per safetensors metadata — embeddings/lm_head/MTP excluded from marketing count) | 2B |
| vision encoder | CogViT, 400M · 24 layers · hidden 1024 · patch 14 · spatial merge 2 | ViT (Qwen2.5-VL family) · patch 14 · spatial merge 2 · deepstack multi-level feature injection |
| connector | lightweight projection + token downsampling → decoder prefix tokens | MLP projection + deepstack into decoder layers |
| text decoder | GLM-0.5B · 16 layers · hidden 1536 · 16 heads / 8 KV (GQA) · head_dim 128 · vocab 59,392 | ~1.5B · 28 layers · hidden 2048 · GQA · vocab ~151,936 |
| positions | M-RoPE-style sections [16, 24, 24] | full 3D M-RoPE (t/h/w) |
| bridge type | identical shape: decoder-only VLM, vision tokens prepended (is_encoder_decoder: false) — no cross-attention on either side | |
| special heads | 1 MTP layer (num_nextn_predict_layers: 1) — drafts ~5.2 tokens/step, vendor claims ~50% throughput gain | none |
| visual token budget | dynamic: 16 → 12,288 tokens/image (12.5k–9.6M px window) | config-bounded megapixels; kaminari pins it to fixed buckets (256–1024 tokens) |
| disk / dtype | one 2.65 GB bf16 safetensors · MIT license | ~4.4 GB bf16 · Apache-2.0 |
| software | transformers ≥ 5.1 (native GlmOcrForConditionalGeneration) · vLLM ≥ 0.19 (MTP spec-decode built in) · SGLang · GGUF (ggml-org, Q4_K_M/Q8_0) · Ollama | transformers · vLLM · kaminari native |
The structural overlap is the headline: same bridge shape, same patch-14 × merge-2 math (28 px/token — kaminari's bucket canvases transfer directly), same GQA-plus-M-RoPE decoder. GLM-OCR drops deepstack (single-level prefix features) and adds an MTP draft head kaminari's core doesn't currently use.
11.2 · the benchmark record
| benchmark | GLM-OCR | context |
|---|---|---|
| OmniDocBench v1.5 (overall) | 94.35 official leaderboard / 94.62 vendor — #1 | PaddleOCR-VL-1.5 94.50 · MinerU2.5 90.93 · DeepSeek-OCR-2 89.17 · dots.ocr 88.41 · Qwen3-VL-235B 89.15 |
| OCRBench (text) | 94.0 | dots.ocr 92.1 · Gemini-3-Pro 91.9 · GPT-5.2 83.7 · PaddleOCR-VL-1.5 75.3 |
| subscores (official) | text edit 0.045 · formula CDM 93.65 · table TEDS 93.89 · TEDS-S 96.50 | PaddleOCR-VL-1.5 wins text edit (0.035) and formula (94.21) |
| real-world (vendor) | receipt KIE 94.5 · seal 90.5 · handwritten 87.0 · code docs 84.7 · multilingual 69.3 | seal recognition is a standout vs dots.ocr 63.0 / Paddle 42.2 |
| olmOCR-bench (self-run) | 75.2 overall (old_scans 37.6 is the weak spot) | from the repo's own .eval_results/ |
| speed (vendor) | 1.86 pages/s PDF · 0.67 img/s, single-replica | hardware unspecified; measured via the two-stage SDK with 32 parallel region workers — not a single-shot decode number |
11.3 · what a port would take
Mapped onto kaminari's actual modules — the port is mostly mechanical because the bridge shape matches:
| kaminari module | work | effort |
|---|---|---|
model/glmocr.py (new) | reuse GlmOcrVisionModel from transformers ≥ 5.1; native 16-layer text decoder port following the qwen3vl.py pattern — fused QKV (16Q/8KV) and gate-up remap apply as-is; M-RoPE sections [16,24,24] slot into the existing Mrope module; MTP layer loaded but inert unless speculative decode is revived | 2–3 focused sessions |
model/cache.py | zero changes — StaticKVCache is parameterized by layers/heads/dims; ~65 KB/token KV at this config | none |
engine/frontend.py | swap processor path (Glm46VProcessor, pop token_type_ids), render the GLM template ([gMASK]<sop>, <|begin_of_image|>…<|end_of_image|>), re-derive fixed buckets (28 px/token math identical); must cap the pixel window — a 200-DPI A4 wants ≈4.9k visual tokens | 1 session |
scripts/quantize_int4.py | retarget the module regex to GLM text layers; no official int4 exists, so the torchao route needs the same pack-verify-smoke treatment Qwen got (a Q4_K_M GGUF exists for the llama.cpp alternative) | 1 session |
engine/runtime.py (optional) | the MTP payoff (+~50% claimed) needs the speculative decode loop that's parked on archive/qwen3-mtp — draft/verify changes the decode contract the compiled one-token step is built around | its own project |
tests/ | tiny glm fixture + parity tests vs transformers forward | 1 session |
Roughly a focused week for a bf16 model-only port with green tests — call it the cleanest second-model port kaminari could ask for — plus open-ended work if the MTP path or the two-stage layout pipeline gets pulled in.
pros of swapping
- purpose-built for kaminari's main use case — #1 on OmniDocBench v1.5, OCRBench text 94.0; receipts/seals/handwriting are first-class training targets, not emergent skills
- a 0.5B decoder is a decode-speed gift — roughly a third of Qwen3-VL-2B's decoder bytes at bf16, vocab 59k vs 152k shrinks the final GEMM; kaminari's int4 logic would compound on top
- 2.65 GB bf16 fits 6 GB with real headroom; GGUF Q4 (~1.6 GB) even more so
- MTP head trained at the weight level — +~50% throughput claim, and the parked
archive/qwen3-mtpspeculative machinery becomes directly relevant - MIT license, first-class transformers/vLLM/SGLang/GGUF support
- fixed prompts = fixed serving contracts —
Text/Table/Formula Recognition:plus JSON-schema KIE map cleanly onto dedicated endpoints - port risk is low — same bridge math kaminari already implements (patch-14 × merge-2, GQA, M-RoPE sections)
cons of swapping
- it is not a general VLM — three fixed recognition prompts + schema-KIE only; kaminari's captioning/querying/VQA scope would need Qwen3-VL kept around anyway (two model paths = two maintenance burdens)
- headline quality rides on the two-stage pipeline — model-only parsing on complex layouts hallucinates/repeats per the tech report; matching the SDK means integrating PP-DocLayoutV3 (Apache-2.0) and region-parallel decoding
- no head-to-head vs Qwen3-VL-2B — the 2B model isn't on OmniDocBench; the swap decision is currently data-free
- visual-token budget is a prefill bomb — up to 12,288 tokens/image (≈4.9k at 200-DPI A4); kaminari's fixed-bucket compile shapes must be re-derived with much larger budgets or aggressively capped; vLLM currently errors above ~3.7k image tokens (issue #106)
- no official int4 — the torchao sidecar route is unproven on
glm_ocrmodules - MTP vs the compiled decode contract — draft/verify doesn't slot into the static-cache one-token loop without real design work; vendor speed numbers don't transfer (unspecified hardware, SDK pipeline)
11.4 · verdict
Keep Qwen3-VL as the engine's general backend; treat GLM-OCR as the specialized OCR track, and let data decide the swap. Concretely:
- today, for OCR output quality: try GLM-OCR through vLLM ≥ 0.19 with its built-in
MTP speculative decoding (
--speculative-config '{"method": "mtp", "num_speculative_tokens": 1}') or the official SDK for full-page parsing — kaminari's parked May experiment (scripts/vllm_image_benchmark.py, narrowed tozai-org/glm-ocrwith an MTP speculative config, perlogs.md) already sketched exactly this path - for the engine: GLM-OCR is the best-justified second
model/package — the bridge matches, the cache/attention/M-RoPE machinery is reusable, and the 0.5B decoder flatters every optimization kaminari already ships — but only after the head-to-head - the deciding experiment: run OmniDocBench-lite (or a fixed page set) through
kaminari's
scripts/ocr_eval.pywith Qwen3-VL-2B bf16/int4 vs GLM-OCR model-only, on the same 3050, same visual-token budget. Quality + tok/s + VRAM on one table; then the swap is an engineering decision, not a vibes decision
12Source trail
| claim | source |
|---|---|
| architecture, state model, decode contract | src/kaminari/engine/*.py, src/kaminari/model/*.py @ main |
| benchmark table | README.md; run JSONs in benchmarks_3vl/ |
| scheduler + endpoints | src/kaminari/engine/scheduler.py, src/kaminari/api.py |
| test suite | tests/ — 27 passed, CPU-only, torch 2.13.0+cpu, transformers 5.16.1 |
| int4 packer/loader | scripts/quantize_int4.py, src/kaminari/model/utils.py |
| repo history (parked branches, logs) | git tags archive/*, logs.md, docs/project.md |
| GLM-OCR facts | huggingface.co/zai-org/GLM-OCR + transformers docs (linked inline in §11) |