How it works
Bootscry is organized around drivers that declare their observation and action capabilities. The working Comet prototype combines captured frames, local OCR, and USB keyboard and mouse input. Serial, camera, and independent input channels are on the roadmap; the shared safety layer is still being built.
Capability-based channels
A video driver can yield frames, a serial driver can yield exact text, and a camera driver could yield images, clips, or audio. Input drivers declare keyboard, mouse, or other actions separately. The interface should require a fresh observation before consequential input; an HID-only driver must be paired with feedback. These are design goals beyond the present Comet prototype.
The intended design uses local OCR first for image channels, with a local vision model planned as a fallback. Serial text would not need OCR. The channel roadmap separates working hardware from lab setup and planned drivers.
The console control loop
Connect to the Comet
bootscry connects to the network KVM over HTTPS and WebSocket. It pins self-signed TLS certificates to prevent credential leakage and handles hardware streamer sleep states transparently.
Read the screen locally
Captured frames can be prepared for Tesseract and RapidOCR. On the measured BIOS frames, Tesseract took 1081 to 1157 ms per frame. OCR supplies word boxes for locating text and clicks; timing varies by screen and host.
Evaluate an exact reader
In planned high-assurance verification, an offline local vision language model (qwen3-vl 30B-A3B evaluated in benchmarks) can serve as an exact reader when verbatim accuracy is critical, scoring 31/32 on strict phrases without character substitutions.
Send input
Keystrokes and mouse coordinates are translated into standard USB HID reports. In kernel tests, pointer positioning lands within 1/32767 of requested coordinates, and typing achieves 25/25 exact strings at full speed.
Coordinate agents, planned
We plan to use machine leases and audit trails to coordinate multiple agents around one physical console. echo.cc is related work on agent coordination, separate from the measured KVM prototype.
Deterministic OCR vs. Vision Models
On a small set of Comet frames, OCR returned coordinates and a local vision model reproduced more phrases verbatim. The vision model is an offline benchmark; a runtime fallback is planned.
| Capability | Local OCR (Tesseract / RapidOCR) | Local VLM (qwen3-vl 30B-A3B) |
|---|---|---|
| Latency on measured BIOS frames | Tesseract 1081 to 1157 ms on measured BIOS frames | 6 to 26 s |
| Bounding boxes for clicks | Word boxes available for click targeting | No coordinate boxes returned |
| Strict verbatim phrases | 22/32 (Tesseract), 16/32 (RapidOCR) | 31/32 verbatim match |
| Role in boot menu catching | Can read a captured result after the hotkey loop | Offline evaluation only; it did not control the measured loop |
The design rule
Hotkeys are sent on a fixed schedule through POST. OCR is useful for reading the resulting screen and finding text after it is stable. A future verification pipeline may compare OCR and a local vision model against an expected phrase.
Boot-Phase Awareness & Hotkey Loops
Catching a boot menu or BIOS setup screen requires sending the right key while firmware accepts input. The current tool runs a timed key loop:
- Trigger a reboot. The earlier benchmark used Ctrl+Alt+Del; the recorded demonstration used a graceful SSH reboot.
- Send the selected hotkey repeatedly for a fixed interval. In the recorded demonstration, Delete taps started when the SSH reboot command returned.
- Capture the resulting screen and verify whether the menu was reached. OCR did not stop the measured key loop.
On the Comet test bench, two documented runs reached AMI Aptio setup. This is evidence on one board, not a success rate across machines.