# Vision model against OCR

**Claim.** On strict verbatim phrases, a local vision model reproduced 31 of 32, Tesseract 22, RapidOCR 16.

**Method.** `benchmark_vlm.py` on real frames from the target: 3 BIOS screens and the console line, 32 phrases. A phrase counts only if reproduced verbatim (whitespace normalised, case kept). The model was qwen3-vl 30B-A3B through Ollama at temperature 0. Tesseract 5.5.3, RapidOCR 3.9.2.

**Sample size.** 4 frames, 32 phrases, 1 vision model. This is small.

**Result.**

| Engine | Verbatim phrases | Seconds per frame |
|---|---|---|
| Tesseract 5.5.3 | 22/32 | about 1 |
| RapidOCR 3.9.2 | 16/32 | about 1.3 |
| qwen3-vl 30B-A3B | 31/32 | 6 to 26 |

Repeat runs of the model gave identical text (3 runs on one frame, 2 on each of four). It read the slashed zeros correctly, which both OCR engines get wrong. Its one miss was turning `@` into `0`. It returns no boxes, so it cannot place a click.

**Limits.**
- Four frames, one model, one screen family, one machine. Do not generalise.
- An earlier run on a long manual page saw the model abridge its output, so completeness depends on the content.
- Published work (FaithC4, arXiv 2607.21617, July 2026) reports vision models rewriting scrambled words on prose pages; that is about prose, not screens like these, and was read by a research agent, not reproduced here.
- Other models (PaddleOCR-VL, DeepSeek-OCR, dots.ocr) are not yet tested.

**Design consequence.** OCR for speed, boxes and waits; a vision model as an exact reader for a screen that matters; a phrase list or regular expression decides on disagreement.

**Raw data.** `docs/evidence-sources/20260929-llm-kvm-landscape.md` (section "What we measured") and `scripts/kvm/benchmark_vlm.py` in this repository. Frames are not published.
