lab

kaminari — a single-GPU Qwen3-VL inference engine, explained

A full walkthrough of kaminari (雷, "thunder"): a minimal, custom inference engine for Qwen3-VL built for OCR/VQA on a 6 GB laptop GPU. What every layer does, how the request flows, where the 3.49× single-request and 9.35× batch speedups come from, how it's tested without a GPU — and what it would take to swap Qwen3-VL for GLM-OCR.

01TL;DR

kaminari is a from-scratch VLM inference engine: a native PyTorch reimplementation of the Qwen3-VL text decoder wrapped around the HF vision tower, a preallocated static KV cache, a torch.compile-shaped one-token decode loop, offline weight-only int4 quantization via torchao, and a FastAPI serving layer with three schedulers (microbatch, continuous slots, SSE streaming). It runs Qwen3-VL-2B at 100.6 tok/s decode — 3.49× the HF generate() baseline — in 2.3 GB of VRAM on an RTX 3050 (6 GB, mobile), and 9.35× aggregate HF throughput at batch 4.

engine

native Qwen3-VL port · fused QKV/gate-up · GQA via SDPA · 3D M-RoPE · deepstack vision features

2,783 lines of production Python

decode

static KV buffers mutated in place · fixed-rank positions · compiled one-token step

100.6 tok/s · 5.18 s / 512-token gen

memory

torchao weight-only int4 on decoder linears · offline sidecar pack · vision + lm_head stay bf16

2.3 GB peak (−45% vs bf16)

serving

FastAPI · microbatch / continuous-slot / streaming schedulers · SSE token stream

batch=2 sweet spot on 6 GB

The whole test suite (27 tests) runs on CPU against a ~1 MB synthetic checkpoint — no GPU needed to verify engine correctness. GPU-only paths are exercised by benchmark scripts on the real box.

02Why kaminari exists

The author runs OCR and local multimodal workloads (receipt transcription, page OCR, captioning, VQA) on a mobile RTX 3050 with 6 GB of VRAM. The stock path — HuggingFace Qwen3VLForConditionalGeneration.generate() — manages 28.8 tok/s on that card, spends memory on dynamic cache growth, and provides no control over batching or preprocessing.

kaminari's design constraints follow from that hardware reality:

  • every MB counts — the model plus a 1024-token KV cache must leave headroom for the 6 GB card's other tenants; int4 weight-only quantization plus a preallocated cache cut peak VRAM from 4.2 GB to 2.3 GB
  • compile shapes must be stable — variable image sizes churn torch.compile guards; fixed visual-token buckets pin the shapes
  • OCR doesn't need sampling — greedy-only decoding removes the sampler from the hot path entirely
  • one GPU owner — a single-GPU runtime doesn't need vLLM's distributed machinery; it needs a tight, auditable loop

The design borrows deliberately from three references: kestrel (moondream's edge inference stack — fixed-shape decode, prefix thinking), gpt-fast (pytorch-native hot path, static KV, weight-only quantization), and nano-vllm (minimal scheduler/paged-cache architecture). kaminari is what falls out of applying those ideas to a multimodal model on a small card.

03Architecture overview

src/kaminari/
├── __init__.py          # exports Engine
├── api.py               # FastAPI server, click CLI, pydantic schemas, SSE
├── vram_utils.py        # VRAM stat helpers
├── engine/
│   ├── types.py         # EngineConfig, Batch, Live, DecodeSlot(+Pool), StreamEvent
│   ├── frontend.py      # cpu preprocessing: bucket → chat template → processor → Batch
│   ├── processor.py     # HF processor loading, dtype + visual-bucket resolution
│   ├── runtime.py       # Engine: generate / generate_stream / prefill / decode / finalize
│   ├── scheduler.py     # MicrobatchScheduler, ContinuousBatchScheduler, StreamingScheduler
│   └── utils.py         # metrics, compiled decode wrapper, sampling, token commit
└── model/
    ├── qwen3vl.py       # Qwen3VLForConditionalGeneration (native) + load_qwen3vl
    ├── modules.py       # TextModel, DecoderLayer, TextAttention (fused QKV), MLP, Mrope
    ├── cache.py         # StaticKVLayer / StaticKVCache (preallocated, slot-addressable)
    └── utils.py         # rope/masks, checkpoint loading, projection fusion, int4 sidecar

Layer responsibilities

layerownsnever touches
api.pytransport: base64/path media → PIL, request schemas, endpoints, server lifecycletensors beyond PIL images
engine/scheduler.pyadmission: bounded queue, coalescing window, slot pool, future scatteringmodel internals
engine/runtime.pyexecution: prefill → decode loop → finalize, live state, metricsHTTP concerns
engine/frontend.pycpu preprocessing: bucketing, chat template, processor invocationcuda
model/*tensors: forward passes, cache, checkpoint loading, fusion, quantization restorerequest concepts

State model

Three immutable dataclasses carry the entire generation lifecycle, updated via dataclasses.replace() — state changes flow through return values, not mutation:

  • Batch — one preprocessed request bundle: request ids, chat texts, padded tensor map from the processor, per-row prompt lengths, images, EOS/pad ids
  • Live — hot decode state: static KV cache reference, accumulated generated_token_ids, next_token_id column, 3D M-RoPE position_ids, runtime cache_position, per-row finished flags
  • DecodeSlot — a Live pinned to a physical cache row for the continuous scheduler; slots live in a fixed DecodeSlotPool sized max_batch_size

The one sanctioned exception to immutability: the KV cache buffers themselves mutate in place — that's the point of a static cache.

04Request lifecycle

blocking path (/generate)

client → /generate→ scheduler: bounded queue + wait_ms coalescing→ microbatch forms (≤ max_batch_size)→ Engine.generate() in a worker thread→ futures scattered back per row

Inside Engine.generate(), five stages run in order:

1 · preprocess (cpu)

Each image is letterboxed onto a fixed visual-token bucket canvas: the frontend picks the aspect-closest divisor grid for the token budget (e.g. 512 tokens → 16×32 → 896×448 px at 28 px/token for a wide image, 448×896 for a tall one), pastes the LANCZOS-resized image centered on a black canvas, then renders the chat template and calls the HF processor with padding="max_length" / truncation=True bounded by kv_cache_max_len − max_new_tokens. The result is one Batch whose padded tensors are trimmed to the batch max prompt length before anything moves to CUDA — pad tokens must never reach the cache.

2 · prefill (one padded batch forward)

Trimmed inputs move to device; the native get_rope_index computes 3D M-RoPE positions and rope_deltas on CPU (integer extraction stays off the GPU path); the static cache is built; a single forward runs vision encode + text prefill with logits_to_keep sliced to each row's last non-pad prompt token. The first token is sampled with greedy argmax immediately — TTFT is the prefill, there is no separate "first decode step".

3 · decode (compiled one-token loop)

Each step calls a torch.compile(mode="reduce-overhead", fullgraph=True)-wrapped closure with fixed-rank tensors only: token ids (batch, 1), positions (3, batch, 1), cache_position (1,), and the full preallocated KV buffers. Greedy selection argmaxes the final logits row outside the compiled callable and stays on-device until commit. Per-row EOS flips a finished flag; finished rows are padded with pad ids so the batch shape never changes mid-generation. The loop runs until every row is finished or max_new_tokens is hit.

4 · finalize

commit_tokens detaches GPU tensors → CPU → python lists, trimming each row at its first EOS; decode_text batch-decodes through the tokenizer. Metrics (preprocessing, prefill, decode, forward, sampling, finalize, TTFT, tok/s, img/s, VRAM) are assembled into EngineMetrics and attached to every row.

streaming path (/generate_stream)

The streaming endpoint bypasses batching entirely: a StreamingScheduler runs Engine.generate_stream() — a synchronous generator sharing the same prefill, then yielding one StreamEvent per decode step — in a worker thread, bridging events to the event loop through an asyncio.Queue. The API layer formats each event as data: {json}\n\n (SSE), checks request.is_disconnected() every cycle, and emits a final done event carrying the full text and metrics.

05The native model port

kaminari does not call HuggingFace's forward(). The text decoder is a native reimplementation in model/; the vision tower is reused from transformers. The split is deliberate: the decoder is where every decode step is spent and where compile/quantization control is needed; the vision tower runs once per request during prefill and HF's implementation is already fused and correct.

componentimplementationwhy
vision towerHF Qwen2_5_VL-family vision encoder, called from the native wrapperruns once per request; deepstack multi-level features and window attention are intricate and not the hot path
text decodernative TextModel: 28 layers, GQA, QK-norm, M-RoPE, SwiGLU MLPfull control of shapes, cache writes, fusion, quantization targets
attentionfused qkv_proj + torch.nn.functional.scaled_dot_product_attention with native GQAone GEMM instead of three per layer; SDPA's GQA path avoids explicit KV replication
MLPfused gate_up_proj → SiLU → elementwise → down_projhalves projection launches; compile-friendly layout
projection fusioncheckpoint remap at load time (fuse_checkpoint_projections): Q/K/V → QKV, gate/up → gate-upthe checkpoint stays upstream-compatible; fusion is a load-time transform, not a fork
positions3D M-RoPE (temporal/height/width), computed per request via get_rope_index, cached as rope_deltas for decodematches Qwen3-VL's multimodal position semantics exactly
KV cacheStaticKVCache: preallocated (max_batch, heads, max_seq, head_dim) buffers per layer, written via cache_position, read as active-prefix slices or full buffers in decodeno per-token torch.cat; stable shapes for compile; slot-addressable rows for continuous batching
the decode contract The compiled decode callable takes exactly five arguments — ids (batch,1), the cache, positions (3,batch,1), cache_position (1,), and a (usually None) mask — and returns one logits row. Everything else (mask construction from the static cache length, cache writes, position increment) happens inside the native forward with no python .item() reductions and no dynamic shapes. That contract is what makes reduce-overhead + fullgraph=True hold across every step of every request.

06The optimization ladder

Single-request decode throughput on the RTX 3050 (Qwen3-VL-2B, 512-token generations, 1024-token KV budget, greedy). Each row is a measurable milestone, kept in benchmarks_3vl/ with the exact run JSON:

optimizationtok/slatencypeak vramvs hf
HF generate() baseline28.8217.77 s—baseline
kaminari HF runtime (early wrapper)25.3820.40 s4173 MB−12.0%
native impl + GQA + fused QKV30.3116.98 s4248 MB+5.2%
custom int4 decoder46.1511.29 s2275 MB+60.2%
torchao int451.889.96 s2275 MB+80.0%
int4 + torch.compile + static KV100.605.18 s2275 MB+249.1%
batch infer, batch=4 (aggregate)269.407.91 s—+834.8%

what each step actually bought

  • native port (+5%) — the win wasn't raw speed; it was control. Owning the forward unlocked the fused projections, the static cache, and compile. (The early HF-wrapper stage was honestly slower than stock — the table keeps that negative result.)
  • int4 (+60→80%) — decode on a 3050 is memory-bandwidth-bound; halving weight bytes nearly doubles decode. VRAM dropped 45% in the same move.
  • compile + static KV (+94% on top of int4) — the big one. Fixed shapes + reduce-overhead (CUDA-graph-backed) removed per-step kernel-launch and python overhead; static buffers removed cache-growth allocations. Together they took 5.18 s per 512-token generation.
  • batching (+835% aggregate) — one padded prefill and one decode step serve N rows; aggregate throughput scales until the card goes VRAM-bound.

the concurrency curve

Concurrent /generate traffic with microbatching enabled surfaces the 3050's limits honestly. At 512 output tokens per request:

batchimages/sectok/smean latency
10.1371.307.50 s
20.24131.218.42 s
40.22121.3818.05 s

batch=2 is the sweet spot: batch=4 doubles latency and loses throughput because operations become VRAM-bound. The merged EngineConfig ships max_batch_size=2 as its default for exactly this reason.

07int4 sidecar quantization

kaminari treats quantization as an offline packaging step, not a runtime option: scripts/quantize_int4.py loads the bf16 model on a CUDA machine, converts only the language-model attention/MLP linears to torchao weight-only int4, and writes a sidecar pack — the non-quantized tensors (vision tower, embeddings, lm_head, norms) stay in a pruned model.safetensors, the packed decoder weights go to qwen3-vl-int4-pack.pt.

why this shape

  • selective targets — only decoder linears are quantized; the vision encoder and lm_head stay bf16, so visual fidelity and the final logits projection keep full precision. The packer validates that no module outside model.language_model.layers.* attention/MLP projections was converted
  • low-memory restore — the engine loads the pruned fp weights, moves to CUDA, then installs the int4 wrappers and reads the sidecar; torchao's tile-packed int4 requires the module to already live on the target device, a subtle ordering constraint the loader handles (and tests pin)
  • explicit, not implicit — quantization="int4" fails hard unless int4_pack_path is provided; there is no silent fallback path that quantizes at load time
open question The memory/speed trade-off is measured; the quality trade-off isn't yet. scripts/ocr_eval.py exists for exactly this — transcribe a fixed image set at bf16 vs int4 and score normalized edit distance against reference texts — but the GPU runs haven't happened. It's the one experiment that would close the loop on the quantization story.

08The serving layer

endpointmethodbehavior
/healthGETload status (starting/ok/error), config echo
/generatePOSTone request; goes through the scheduler as a batch of 1
/generate_batchPOSTexplicit row batch; each row is a scheduler job so microbatching still decides grouping
/generate_streamPOSTSSE: one data: {json} per token + final done with metrics; disconnect-aware

three schedulers, one engine

  • MicrobatchScheduler (default) — bounded asyncio.Queue (backpressure via max_queue_size), a short coalescing window after the first job so concurrent requests batch together, one Engine.generate() per batch in a worker thread, futures scattered back to rows. Per-job scheduler metrics (queue wait, microbatch wait, batch id/size) ride alongside engine metrics in every response.
  • ContinuousBatchScheduler — admits requests into free DecodeSlots and ticks active slots one token at a time, finalizing each row the moment it finishes instead of waiting for the whole batch. The slot pool owns a shared static cache with per-row physical rows; copy_row_from migrates a finished row's KV state.
  • StreamingScheduler — no batching, no slots: one stream at a time, token events bridged from the engine thread to the event loop.

The engine itself stays synchronous and batch-shaped — all async concerns live in the scheduler layer, and the API layer never touches tensors. Server defaults match the measured sweet spot: int4 on, max_batch_size=2, 50 ms coalescing window, kv_cache_max_len=2048.

09Correctness engineering

An inference engine that can only be tested on a GPU is an engine you can't refactor. kaminari's suite runs entirely on CPU: scripts/create_tiny_qwen3_vl_checkpoint.py synthesizes a ~1 MB Qwen3-VL-compatible checkpoint (tiny-but-structurally-exact config: 2 layers, 4 heads, the real processor and chat template) into tests/fixtures/tiny_qwen3_vl_cpu, and every test runs the real engine path against it.

test filewhat it pins
test_tiny_qwen3_vl_fixture.pyend-to-end: frontend input building, native model forward, batch generation, prefill/decode Live state invariants (cache position, token accumulation, completion predicate)
test_api_server.pyhealth / generate / concurrent microbatch serving against a dummy runtime — transport correctness without weights
test_decode_slots.pystatic KV semantics: 1D position expansion per row, slot-row writes, active-prefix reads, copy_row_from
test_quantization.pyint4 sidecar contracts: pack filename, module targeting, cuda-before-pack ordering, weight filtering
test_runtime_compile.pythe compiled decode wrapper: decode_token closure, reduce-overhead + fullgraph=True
test_utils_perf.pyVRAM stat helpers degrade gracefully on CPU

27 tests, ruff check clean across src/ tests/ scripts/, deps pinned to CPU wheels via [tool.uv.sources] so uv sync + pytest works on any laptop. Testing locally is the whole story — no CI/CD, no remote runners.

10Repo state & roadmap

where it stands

  • main carries everything: native engine, int4 sidecar, three schedulers, SSE streaming, green CPU suite, ruff-clean lint, MIT license, committed benchmark history (benchmarks_3vl/ JSONs back the README table)
  • parked as tags — archive/qwen3-mtp (qwen3.5 adapter + MTP speculative decoding, 15 commits) and archive/decode-slots-v1 (slot exploration, 12 commits); restore with git checkout <tag>
  • documented — docs/project.md maps every module, the request path, the scripts inventory, and the simplification proposal; logs.md carries the full edit history

next unlocks, in the order that makes sense

unlockwhat it iswhy it's next
int4 quality numbersrun scripts/ocr_eval.py bf16-vs-int4 on the OCR bench setcloses the one open scientific question; needs one GPU session
streaming verification + benchGPU-verify /generate_stream, add streaming numbers to the README tablethe streaming path merged untested on real hardware
core collapsethe simplification proposal: strip serving from the core, target 1–1.5k linesthe runtime should own one Live per call, not a request map
paged-kv extractionstandalone block allocator + block table → continuous batching without slot copiessystems-depth follow-up; a second project, not core work
prefix cachingreuse the fixed BOS+image-token prefix across OCR calls to the same page setkestrel-style; large prefill savings for repeated OCR prompts

11Swap analysis: replacing Qwen3-VL with GLM-OCR

GLM-OCR (zai-org, Jan 2026) is a purpose-built document-OCR VLM — and a genuinely interesting swap candidate for kaminari, because kaminari's stated hot path (page/receipt OCR) is exactly what it was trained for. This section is a fact-first look at what the model is, how it compares, and what an engine port would actually involve.

11.1 · what it is

propertyGLM-OCRQwen3-VL-2B (kaminari today)
total params0.9B vendor claim (1.33B per safetensors metadata — embeddings/lm_head/MTP excluded from marketing count)2B
vision encoderCogViT, 400M · 24 layers · hidden 1024 · patch 14 · spatial merge 2ViT (Qwen2.5-VL family) · patch 14 · spatial merge 2 · deepstack multi-level feature injection
connectorlightweight projection + token downsampling → decoder prefix tokensMLP projection + deepstack into decoder layers
text decoderGLM-0.5B · 16 layers · hidden 1536 · 16 heads / 8 KV (GQA) · head_dim 128 · vocab 59,392~1.5B · 28 layers · hidden 2048 · GQA · vocab ~151,936
positionsM-RoPE-style sections [16, 24, 24]full 3D M-RoPE (t/h/w)
bridge typeidentical shape: decoder-only VLM, vision tokens prepended (is_encoder_decoder: false) — no cross-attention on either side
special heads1 MTP layer (num_nextn_predict_layers: 1) — drafts ~5.2 tokens/step, vendor claims ~50% throughput gainnone
visual token budgetdynamic: 16 → 12,288 tokens/image (12.5k–9.6M px window)config-bounded megapixels; kaminari pins it to fixed buckets (256–1024 tokens)
disk / dtypeone 2.65 GB bf16 safetensors · MIT license~4.4 GB bf16 · Apache-2.0
softwaretransformers ≥ 5.1 (native GlmOcrForConditionalGeneration) · vLLM ≥ 0.19 (MTP spec-decode built in) · SGLang · GGUF (ggml-org, Q4_K_M/Q8_0) · Ollamatransformers · vLLM · kaminari native

The structural overlap is the headline: same bridge shape, same patch-14 × merge-2 math (28 px/token — kaminari's bucket canvases transfer directly), same GQA-plus-M-RoPE decoder. GLM-OCR drops deepstack (single-level prefix features) and adds an MTP draft head kaminari's core doesn't currently use.

11.2 · the benchmark record

benchmarkGLM-OCRcontext
OmniDocBench v1.5 (overall)94.35 official leaderboard / 94.62 vendor — #1PaddleOCR-VL-1.5 94.50 · MinerU2.5 90.93 · DeepSeek-OCR-2 89.17 · dots.ocr 88.41 · Qwen3-VL-235B 89.15
OCRBench (text)94.0dots.ocr 92.1 · Gemini-3-Pro 91.9 · GPT-5.2 83.7 · PaddleOCR-VL-1.5 75.3
subscores (official)text edit 0.045 · formula CDM 93.65 · table TEDS 93.89 · TEDS-S 96.50PaddleOCR-VL-1.5 wins text edit (0.035) and formula (94.21)
real-world (vendor)receipt KIE 94.5 · seal 90.5 · handwritten 87.0 · code docs 84.7 · multilingual 69.3seal recognition is a standout vs dots.ocr 63.0 / Paddle 42.2
olmOCR-bench (self-run)75.2 overall (old_scans 37.6 is the weak spot)from the repo's own .eval_results/
speed (vendor)1.86 pages/s PDF · 0.67 img/s, single-replicahardware unspecified; measured via the two-stage SDK with 32 parallel region workers — not a single-shot decode number
the honest gaps Qwen3-VL-2B is not on the OmniDocBench leaderboard — the 89.15 entry is the 235B model — so “GLM-OCR beats kaminari's current model at document parsing” is suggested, not proven, until someone runs the bench locally. And the vendor speed figures come from the two-stage pipeline (PP-DocLayoutV3 layout model → cropped-region OCR in parallel), whose layout decomposition the tech report itself says small models need on complex pages — model-only parsing on hard layouts hallucinates and repeats.

11.3 · what a port would take

Mapped onto kaminari's actual modules — the port is mostly mechanical because the bridge shape matches:

kaminari moduleworkeffort
model/glmocr.py (new)reuse GlmOcrVisionModel from transformers ≥ 5.1; native 16-layer text decoder port following the qwen3vl.py pattern — fused QKV (16Q/8KV) and gate-up remap apply as-is; M-RoPE sections [16,24,24] slot into the existing Mrope module; MTP layer loaded but inert unless speculative decode is revived2–3 focused sessions
model/cache.pyzero changes — StaticKVCache is parameterized by layers/heads/dims; ~65 KB/token KV at this confignone
engine/frontend.pyswap processor path (Glm46VProcessor, pop token_type_ids), render the GLM template ([gMASK]<sop>, <|begin_of_image|>…<|end_of_image|>), re-derive fixed buckets (28 px/token math identical); must cap the pixel window — a 200-DPI A4 wants ≈4.9k visual tokens1 session
scripts/quantize_int4.pyretarget the module regex to GLM text layers; no official int4 exists, so the torchao route needs the same pack-verify-smoke treatment Qwen got (a Q4_K_M GGUF exists for the llama.cpp alternative)1 session
engine/runtime.py (optional)the MTP payoff (+~50% claimed) needs the speculative decode loop that's parked on archive/qwen3-mtp — draft/verify changes the decode contract the compiled one-token step is built aroundits own project
tests/tiny glm fixture + parity tests vs transformers forward1 session

Roughly a focused week for a bf16 model-only port with green tests — call it the cleanest second-model port kaminari could ask for — plus open-ended work if the MTP path or the two-stage layout pipeline gets pulled in.

pros of swapping

  • purpose-built for kaminari's main use case — #1 on OmniDocBench v1.5, OCRBench text 94.0; receipts/seals/handwriting are first-class training targets, not emergent skills
  • a 0.5B decoder is a decode-speed gift — roughly a third of Qwen3-VL-2B's decoder bytes at bf16, vocab 59k vs 152k shrinks the final GEMM; kaminari's int4 logic would compound on top
  • 2.65 GB bf16 fits 6 GB with real headroom; GGUF Q4 (~1.6 GB) even more so
  • MTP head trained at the weight level — +~50% throughput claim, and the parked archive/qwen3-mtp speculative machinery becomes directly relevant
  • MIT license, first-class transformers/vLLM/SGLang/GGUF support
  • fixed prompts = fixed serving contracts — Text/Table/Formula Recognition: plus JSON-schema KIE map cleanly onto dedicated endpoints
  • port risk is low — same bridge math kaminari already implements (patch-14 × merge-2, GQA, M-RoPE sections)

cons of swapping

  • it is not a general VLM — three fixed recognition prompts + schema-KIE only; kaminari's captioning/querying/VQA scope would need Qwen3-VL kept around anyway (two model paths = two maintenance burdens)
  • headline quality rides on the two-stage pipeline — model-only parsing on complex layouts hallucinates/repeats per the tech report; matching the SDK means integrating PP-DocLayoutV3 (Apache-2.0) and region-parallel decoding
  • no head-to-head vs Qwen3-VL-2B — the 2B model isn't on OmniDocBench; the swap decision is currently data-free
  • visual-token budget is a prefill bomb — up to 12,288 tokens/image (≈4.9k at 200-DPI A4); kaminari's fixed-bucket compile shapes must be re-derived with much larger budgets or aggressively capped; vLLM currently errors above ~3.7k image tokens (issue #106)
  • no official int4 — the torchao sidecar route is unproven on glm_ocr modules
  • MTP vs the compiled decode contract — draft/verify doesn't slot into the static-cache one-token loop without real design work; vendor speed numbers don't transfer (unspecified hardware, SDK pipeline)

11.4 · verdict

Keep Qwen3-VL as the engine's general backend; treat GLM-OCR as the specialized OCR track, and let data decide the swap. Concretely:

  • today, for OCR output quality: try GLM-OCR through vLLM ≥ 0.19 with its built-in MTP speculative decoding (--speculative-config '{"method": "mtp", "num_speculative_tokens": 1}') or the official SDK for full-page parsing — kaminari's parked May experiment (scripts/vllm_image_benchmark.py, narrowed to zai-org/glm-ocr with an MTP speculative config, per logs.md) already sketched exactly this path
  • for the engine: GLM-OCR is the best-justified second model/ package — the bridge matches, the cache/attention/M-RoPE machinery is reusable, and the 0.5B decoder flatters every optimization kaminari already ships — but only after the head-to-head
  • the deciding experiment: run OmniDocBench-lite (or a fixed page set) through kaminari's scripts/ocr_eval.py with Qwen3-VL-2B bf16/int4 vs GLM-OCR model-only, on the same 3050, same visual-token budget. Quality + tok/s + VRAM on one table; then the swap is an engineering decision, not a vibes decision

12Source trail

claimsource
architecture, state model, decode contractsrc/kaminari/engine/*.py, src/kaminari/model/*.py @ main
benchmark tableREADME.md; run JSONs in benchmarks_3vl/
scheduler + endpointssrc/kaminari/engine/scheduler.py, src/kaminari/api.py
test suitetests/ — 27 passed, CPU-only, torch 2.13.0+cpu, transformers 5.16.1
int4 packer/loaderscripts/quantize_int4.py, src/kaminari/model/utils.py
repo history (parked branches, logs)git tags archive/*, logs.md, docs/project.md
GLM-OCR factshuggingface.co/zai-org/GLM-OCR + transformers docs (linked inline in §11)