Vision model against OCR
Claim. On strict verbatim phrases, a local vision model reproduced 31 of 32, Tesseract 22, RapidOCR 16.
Method. benchmark_vlm.py on real frames from the target: 3 BIOS screens and the console line, 32 phrases. A phrase counts only if reproduced verbatim (whitespace normalised, case kept). The model was qwen3-vl 30B-A3B through Ollama at temperature 0. Tesseract 5.5.3, RapidOCR 3.9.2.
Sample size. 4 frames, 32 phrases, 1 vision model. This is small.
Result.
| Engine | Verbatim phrases | Seconds per frame |
|---|---|---|
| Tesseract 5.5.3 | 22/32 | about 1 |
| RapidOCR 3.9.2 | 16/32 | about 1.3 |
| qwen3-vl 30B-A3B | 31/32 | 6 to 26 |
Repeat runs of the model gave identical text (3 runs on one frame, 2 on each of four). It read the slashed zeros correctly, which both OCR engines get wrong. Its one miss was turning @ into 0. It returns no boxes, so it cannot place a click.
Limits.
- Four frames, one model, one screen family, one machine. Do not generalise.
- An earlier run on a long manual page saw the model abridge its output, so completeness depends on the content.
- Published work (FaithC4, arXiv 2607.21617, July 2026) reports vision models rewriting scrambled words on prose pages; that is about prose, not screens like these, and was read by a research agent, not reproduced here.
- Other models (PaddleOCR-VL, DeepSeek-OCR, dots.ocr) are not yet tested.
Design consequence. OCR for speed, boxes and waits; a vision model as an exact reader for a screen that matters; a phrase list or regular expression decides on disagreement.
Raw data. docs/evidence-sources/20260929-llm-kvm-landscape.md (section "What we measured") and scripts/kvm/benchmark_vlm.py in this repository. Frames are not published.