NAS boot loop recovered through the Comet with local OCR
Claim
A GL.iNet Comet KVM, with local OCR running on a separate machine, was used to diagnose and fix a real boot loop on a headless NAS, with no screen or keyboard on it. This is one incident on one machine. It is not a success rate.
What happened
Everything in this section comes from the field manual named at the bottom. Numbers are as the manual states them.
- The target was a UGREEN DXP4800 Plus NAS on UGOS Pro (a Debian 12 base). After a system update and reboot it left the network.
- SSH refused connections and the web ports were dark. The machine rebooted every 3 to 4 minutes, and the manual says the reboot came every 180 seconds.
- When a frame was caught, the machine sat at a systemd emergency-mode prompt. Normal recovery access did not work.
- A hardware watchdog on the board reset the machine after exactly 180 seconds, so any repair had a 3-minute window.
- The BIOS boot window was under 400 milliseconds. Plain F12 was ignored. The board needed Ctrl+F12 for the boot menu and Ctrl+F2 for setup.
- The system used a read-only base image with an overlay. Edits to the base image did not persist until the overlay was changed.
- Root cause: a log-restore service tried to copy far more saved journal data into a small RAM-backed log area than it could hold. The boot dependency failed, the system entered emergency mode, and the watchdog restarted the loop.
What the KVM and OCR did
- The Comet supplied the screen over HDMI and keyboard input over USB. The manual also describes ATX wiring, but the later reference setup was verified without an ATX board. This page does not claim a verified KVM power-button action.
- The manual says the Comet's API was worked out from network inspection and firmware analysis: session login, a WebSocket for keys and mouse, and a snapshot endpoint for frames.
- Frames were transient, so snapshots were sent to a second machine for local OCR. The manual reports under 600 ms per frame on its console frames; the later BIOS benchmark measured different times.
- A script rebooted the target, watched the video signal drop, then pressed Ctrl+F12 every 150 ms for 14 seconds, reading the screen until the boot menu of the attached recovery USB appeared.
- Typing into the emergency shell first failed because Windows line endings (
\r\n) turned into extra Enter presses. The fix was to strip them and send commands in base64 chunks. - Image preparation before OCR (2x upscale, inversion of dark screens, contrast, binarization) raised character accuracy on the manual's console test frames from about 82% to 100%.
- The manual also records a false lead: the KVM lost video and keyboard because a USB drive was plugged into the port next to the KVM cable and the HDMI line came unseated. Checking kernel logs, USB topology and EDID showed the cause was physical, not a GPU driver fault.
What fixed it
From a live USB system the overlay was mounted and edited. The manual lists fixes including:
- Tolerating unavailable external mounts during boot.
- Changing log restoration so archived data could not exhaust the RAM-backed log area or block startup.
- Adding a recurring boot-health check and preserving a KVM recovery path.
The manual's closing health check reports 21 passed, 2 warnings and 0 failures across 23 checks. The 2 warnings are the KVM's USB link and HDMI signal, which were offline at that moment.
Limits
- One incident, one NAS, one KVM. No repeat runs and no success rate.
- The manual is a post-mortem written by the operator's own team. There is no independent log of the run, and the manual does not say which steps were typed by an agent and which by a person. That the run was unattended is the operator's account only.
- Physical work was needed: the manual describes a drive being plugged in and cables being disturbed at the target.
- The OCR engine comparison in the manual (Tesseract, RapidOCR on two hosts) uses a small set of frames and is separate from the measurements on the other evidence pages.
- Nothing here is a claim about any other vendor's KVM.
Operator's account, not in the manual
These statements come from the operator, who confirmed them in chat on 2026-09-29. The manual does not document them and there is no log of them here. In the operator's account:
- The trigger was a routine upgrade of the NAS through its vendor's web interface. The operator waited for over an hour, then attached a screen to the NAS and saw it "frozen". Nothing on the network showed why. That is what led to trying an agent with a KVM. This fits the manual, whose section 2 says the host disappeared after a system update and reboot; the manual does not record the web-interface detail or the hour of waiting.
- The AI work began on a Windows workstation. The Comet was tried there first, to see whether an agent could read the screen.
- The Comet was then given typing exercises, and extended step by step.
- The operator reports leaving the agent working overnight and finding the NAS recovered in the morning. The post-mortem records the repair but not a step-by-step action log, so it cannot establish which steps the agent took on its own.
- The work then continued on a Linux server. Later the Comet was connected to a dedicated Linux test machine for repeatable measurements.
The manual itself says the recovery commands were sent "from the Windows management workstation" (Discovery 4), which fits this account. It does not name the machine.
Source material
The operator's full incident manual and recovery runbook are held in the project repository. This public summary omits host access and account configuration details. Contact the maintainer for a review copy of the relevant redacted records.