Known Issues
Safety & Security: Data Collection Controls
See ../security.md for data collection controls, privacy settings, and the full reference including set_config options, watcher management, and noise cleanup.
v0.2.7 — Statistical priors, noise cleanup, config gating
New in v0.2.7:
| Feature | Status | Notes |
|---|---|---|
Statistical prior model (priors_stats) |
Code complete | Element frequency + position per-window, updated on every training save. Go-native, no Python. |
| Prior confidence adjustment | Code complete | Applied in ONNXDetect after NMS. Boosts expected elements, suppresses outliers. Gated by prior_adjustment config. |
export_yolo_dataset |
Code complete | Exports unused samples as YOLO-format dataset for external training. |
training_cleanup_noise |
Code complete | Deletes signal_level=0 samples older than N hours. Supports dry_run. |
training_enabled config |
Code complete | When false, disables all auto-save snapshots (actions + watcher). Default: true. |
| Element priors not yet verified with real accumulated data | Untested | Priors are updated asynchronously; first detections in a session have no priors loaded until the first loadPriorsFromDB call. |
| Done v0.2.37 | retention_days config auto-prunes samples older than N days. Background pruner runs every 6 hours. |
Test Session: v0.1.2 — 2026-06-27
Tools Tested & Working (22)
| Tool | Status | Notes |
|---|---|---|
get_screen_size |
✅ | 3200×900 (virtual — 2×1600) |
get_system_info |
✅ | Hostname, OS, RAM |
get_cursor_position |
✅ | Confirms second screen (x=2142) |
get_battery |
✅ | “no_battery” (desktop) |
get_disk_usage |
✅ | 6 drives enumerated |
get_idle_time |
✅ | |
get_keyboard_layout |
✅ | 00000409 (en-US) |
get_network_info |
✅ | 8 IPs, hostname |
get_screen_dpi |
✅ | 2 monitors, both 96dpi 100% |
get_uptime |
✅ | |
get_volume |
✅ | 49% |
list_displays |
✅ | DISPLAY1 + DISPLAY2 (v0.1.8) |
list_windows |
✅ | 24 windows |
list_processes |
✅ | 195 processes |
get_active_window |
✅ | OpenCode |
get_pixel_color |
✅ | #2b2a33 |
get_clipboard |
✅ | “ready” |
find_window |
✅ | OpenCode handle 463694 |
get_window_state |
✅ | Visible, maximized, rect (1592,-8,3208,860) |
get_display_modes |
✅ | 37 modes for DISPLAY1 |
screenshot |
✅ | Returns base64 PNG |
find_text_and_click |
✅ | Found “OpenCode” and clicked |
Tools Tested This Session (additional)
| Tool | Status | Notes |
|---|---|---|
set_clipboard |
✅ | Write “test from go-mcp-computer-use” |
open_url |
✅ | https://example.com — opens in default browser |
scroll |
✅ | -3 clicks |
click (left) |
✅ | 100,100 |
click (double) |
✅ | 200,200 clicks=2 |
click (right) |
✅ | 300,300 button=right |
click (middle) |
✅ | 400,400 button=middle |
drag |
✅ | (500,500) → (600,600) |
uia_find |
✅ | GitHub Desktop window found |
wait |
✅ | v0.1.5 — B6: Wait() calc was 1Mx too long |
hover |
✅ | v0.1.5 — B6: same root cause as wait |
screenshot_element |
✅ | v0.1.3 — B5: clamps to screen bounds |
uia_get_text |
✅ | v0.1.7 — B4: nil pattern check added |
Tools Not Yet Tested (3)
| Tool | Reason |
|---|---|
type_text |
Interactive — needs terminal |
key_press |
Interactive |
key_sequence |
Interactive |
Test Session: v0.2.27 — 2026-06-30
Tools Tested (v0.2.27 additions)
| Tool | Status | Notes |
|---|---|---|
find_all_images |
✅ | Degenerate template (1×1 constant color) → NCC skipped → ONNX detected 68+ COCO objects → OCR 200+ words. Valid template at 0.99 threshold → same cascade. |
find_image |
✅ | Degenerate template → ONNX best element returned {x:-19, y:4, w:168, h:47, score:1}. |
click (button=middle) |
✅ | No error on middle click at (100,100). |
scroll (horizontal=true) |
✅ | No error with 3 clicks horizontal scroll. |
get_window_state (fullscreen) |
✅ | Returns fullscreen: false for normal windows, true for borderless fullscreen games. |
ocr_languages |
✅ | Native COM — 2 languages (en-GB, en-US). PowerShell fallback eliminated. |
Bug Reports
B1. get_brightness — parse failure (fixed v0.1.4)
get_brightness — parse failureError: parse brightness: strconv.Atoi: parsing "": invalid syntax
Root cause: PowerShell/WMI brightness query returns empty string instead of a numeric value. Likely because the display (DVI/HDMI desktop monitor) doesn’t support WMI brightness control (laptops only).
Fix: Return a meaningful error or handle gracefully (e.g., "not supported" instead of crash).
B2. list_audio_devices — returns null (fixed v0.1.6)
list_audio_devices — returns nullResult: {"devices":null}
Issue: No audio devices enumerated. PowerShell Get-AudioDevice -List may not be installed on this system (requires AudioDeviceCmdlets module).
Fix: Return empty slice [] instead of nil slice null.
B3. list_displays — second monitor not enumerated (fixed v0.1.8)
list_displays — second monitor not enumeratedEvidence:
- Cursor position: x=2142 (primary is 1600 wide, so x≥1600 = second screen)
- OpenCode window rect:
left=1592, right=3208, width=1616— spans across to a second screen at x~1600
But list_displays only returns DISPLAY1.
Root cause: monitorEnumProc callback gated processing on mi.Flags&1 != 0 (MONITORINFOF_PRIMARY = 0x1). Non-primary monitors were silently skipped.
Fix: Removed the primary-only gate — all enumerated monitors are now included, with Primary: mi.Flags&1 != 0 set correctly per-monitor.
B4. uia_get_text / uia_invoke — server disconnect (fixed v0.1.7)
uia_get_text / uia_invoke — server disconnectAction:
uia_get_text(name="Taskbar")— connection lostuia_get_text(name="GitHub Desktop")— connection lostuia_invoke(name="Taskbar")— connection lost Root cause:GetCurrentPatternreturnsS_OKwithnilpointer when element doesn’t support pattern. Code then callscomRelease(0)and vtbl methods on0— nil pointer dereference crashes MCP transport. Fix: Addedp == 0check ingetCurrentPattern()— returns clear error instead of nil dereference.
B5. screenshot_element — negative coordinates rejected (fixed v0.1.3)
screenshot_element — negative coordinates rejectedError: x=-8 out of bounds (screen width=1600)
Context: Firefox window handle 132490 had rect left=-8 (window decorations positioned off-screen by Windows snap behaviour).
Fix: Screenshot element should clamp/clip the region to screen bounds rather than rejecting negative coordinates. A window with off-screen decorations is a normal state (Aero Snap, multi-monitor).
B6. hover — consistently hangs/”Tool execution aborted” (fixed v0.1.5)
hover — consistently hangs/”Tool execution aborted”Root cause: Wait() used int64(duration) * 10000 (where duration is nanoseconds), producing a timeout 1 million times too long. Wait(100ms) blocked for ~27.7 hours.
Fix: Changed to int64(ms) * 10000 (1ms = 10,000 × 100ns intervals). Same fix applies to B7.
B7. wait — “Tool execution aborted” (same as B6, fixed v0.1.5)
wait — “Tool execution aborted”B8. find_text_and_click — steals focus
Observation: Calling find_text_and_click brings the target window to foreground. This is expected behavior for a computer-use tool, but worth documenting as a caveat.
Workaround: None — by design.
B9. Keyboard input blocked by UIPI on elevated windows (fixed v0.1.9)
Observation: All type, type_and_submit, key_press, select_all_and_type return ok but input does not reach elevated (Administrator) PowerShell.
Root cause: Windows UIPI — SendInput with KEYEVENTF_UNICODE from non-elevated MCP server is silently blocked from reaching admin-elevated windows.
Fix: Added isForegroundElevated() check using OpenProcess + GetTokenInformation(TokenElevation). Returns clear warning message instead of silent failure.
B10. click may silently fail on elevated windows (documented)
click may silently fail on elevated windowsNote: Click uses SetCursorPos + SendInput mouse events. Same UIPI restriction applies — no error feedback when targeting admin windows. Unlike keyboard (which always targets foreground window), mouse targets coordinates, making elevation check impractical. Run MCP server elevated to avoid this.
B11. KeyPress modifier key ordering — Ctrl+C sends C before pressing Ctrl
Observation: KeyPress(["CTRL", "C"]) sends sendUnicode('C') first, then presses Ctrl down, then releases Ctrl. The C arrives before Ctrl is held, so Ctrl+C works as just c in most apps/games.
Root cause: KeyPress was splitting keys into three phases: Unicode chars → VK downs → VK ups. Modifiers and their target keys were sent in separate batches, not interleaved.
Fix applied: Replaced batch processing with in-order processing. Modifiers are pressed immediately when encountered, then their paired keys are sent while the modifier is held. All pressed modifiers are released in reverse order after the key sequence.
B12. KEYEVENTF_UNICODE may not work in game engines (fixed v0.2.8)
KEYEVENTF_UNICODE may not work in game enginesObservation: sendUnicode injects characters via KEYEVENTF_UNICODE, which synthesizes WM_CHAR messages. Game engines using raw input or GetAsyncKeyState for keyboard polling typically don’t see Unicode-injected characters — they check VK codes and scan codes instead. Same issue affects terminals, code editors, and browser input fields in some configurations.
Root cause: All keyboard input used KEYEVENTF_UNICODE — letters, digits, TypeText, TypeAndSubmit. Only modifier keys and special keys (arrows, F-keys) used VK codes.
Fix: Removed KEYEVENTF_UNICODE entirely. Rewrote keyboard input to use VK codes for everything:
- Letters A-Z/a-z →
VK_A–VK_Z(0x41–0x5A) with Shift state for case - Digits and punctuation → VK codes with Shift mapping for symbols
- TypeText/TypedAndSubmit →
sendCharWithVK()handles shift state per character KeyPressmodifier order fixed: modifiers are pressed before their target keys
Prompt Engineering: Learn-Once-Reuse-Forever Pattern
The MCP server exposes 161 tools, but an AI agent using them starts cold every session — no knowledge of:
- What windows exist on this user’s desktop and where they’re positioned
- How specific applications render (Firefox tab bar vs URL bar, Outlook email list vs reading pane)
- What sequences of tool calls successfully completed a task last time
- What edge cases exist (Firefox containers, UIPI elevation blocks, OCR timing)
The Pattern
After any successful GUI interaction sequence, the AI should:
- Store the sequence as a named macro/recipe (e.g., “open_chrome_search_google”)
- Annotate it with application name, window layout details, and screen dimensions
- Scope it to the application so it’s reusable across sessions
- Next time the same task is asked, recall the stored sequence and execute it directly — no need to rediscover coordinates and timings
Example memory entry
Application: Firefox (v134+ with Multi-Account Containers)
Window size: 1123x791 (positioned at x=295, y=39)
Tab bar: y≈50-70
URL bar: y≈90-110 (click at x=350, y=105 to focus)
Container new-tab: Ctrl+T bypasses popup, or click "No Container" at x=830,y=105
Bookmarks bar: y≈120-140 (when enabled)
Sequence "open_google_and_search":
1. focus_window(handle=132490)
2. click(x=350, y=105) — focus URL bar
3. type_and_submit("google.com")
4. wait(4000)
5. type_and_submit("search query")
6. wait(3000)
7. scroll(clicks=-6)
8. ocr(x=295, y=140, w=1123, h=700) — read results below URL bar
Why this matters
Without this pattern, the AI wastes time and tokens rediscovering basic facts each session — where the URL bar is, that Firefox uses containers, that scroll takes negative values for down. With it, the AI builds a living knowledge base that compounds with every session.
Lessons Learned (from live testing)
L1. Screen layout awareness is critical — always survey before acting
Problem: Commands like click(x,y) / type_text fail silently when the AI doesn’t know the screen layout — what windows exist, their positions, what UI elements are where.
Example from session: Firefox had:
- Window rect:
{left:295, top:39, right:1418, bottom:830}(1123×791) - Multi-Account Containers extension modified the new-tab
+button behavior — clicking it showed a container picker menu instead of opening a blank tab - The URL bar was at y≈96 (below the tab bar at y≈56), not at the very top of the window
- Tab bar labels were partially visible but non-obvious (“< Intern PocketStac”, “discwc”)
Procedure for any GUI interaction:
get_window_state(handle)— get target window positionocr(x,y,w,h)over the window region — see what’s actually displayed- Locate the target element (button, text field, link) from OCR coordinates
- Click at the element’s center position
- Verify with another OCR call after action
Firefox-specific layout (tested v134+):
- Tab bar: y≈50-70 (depends on title bar visibility). Compact tab mode changes spacing.
- URL bar: y≈90-110. Contains: padlock icon + “about:” or URL text.
- Container extensions add a popup menu on
+click — must click “No Container” to open a regular tab. - Bookmarks toolbar: y≈120-140 (if enabled). Can shift content down.
- Window top (y=39 for this session) includes the OS window title bar (if not maximized).
L2. Tools return “ok” even when the action had no visible effect
type, key_press, type_and_submit, click all return ok — but the input may hit the wrong element or be dropped by UIPI. Always verify with OCR/screenshot after each action.
L3. Firefox containers intercept the + new tab button
Firefox Multi-Account Containers changes the new-tab + behavior — instead of opening a blank tab immediately, it shows a popup asking which container to use. Click “No Container” (≈x=830, y=105 in this layout) or use Ctrl+T which bypasses the popup.
B13. ONNX tools require CGO (disabled in default CGO_ENABLED=0 build)
Observation: go build ./... fails with build constraints exclude all Go files for github.com/yalue/onnxruntime_go when CGO_ENABLED=0. The onnxruntime_go library uses cgo for native shared library bindings.
Impact: ONNX ML tools (onnx_detect, onnx_status, onnx_download) are excluded from CGO-free builds. All other tools work — including the Go-native transformer engine (action prediction), which is pure Go via Gorgonia.
Workaround: Build with CGO_ENABLED=1 and a C compiler:
- Zig cc:
CC="zig cc" CGO_ENABLED=1 go build -o mcp-server.exe .\cmd\mcp-server\ - GCC (Mingw-w64):
CGO_ENABLED=1 go build -o mcp-server.exe .\cmd\mcp-server\
Status: Documented in v0.2.x. Not a bug — by design. CGO only needed for ONNX; transformer is pure Go.
B14. ONNX YOLO11n model uses unsupported opset 22
Done v0.2.13 — ONNX Runtime DLL updated to v1.26.0 which supports opset 22. Detection works.
B15. No action verification — tools return “ok” for API call success, not for real-world effect
Severity: High — this is the #1 reason the server feels “hit or miss” to an AI agent.
Every action tool (click, type, key_press, scroll, drag, etc.) uses fire-and-forget Win32 APIs (SendInput, SetCursorPos, keybd_event). They return success if the API call didn’t crash — not if the action had any visible effect on screen. There is no built-in verification loop between executing an action and confirming it achieved its goal.
What tools actually return:
click(x=100, y=200)→"ok"meansSetCursorPos+SendInputreturned non-zero. It does not mean the click hit an interactive element, changed UI state, or reached the intended window.type("hello")→"ok"means key events were queued. It does not mean text appeared in an input field.key_press(["Enter"])→"ok"means the key was sent. It does not mean anything happened on screen.screenshot/ocr→ Returns actual pixel/text data, but the AI has no way to confirm it matches what should be on screen.
The AI’s blind spots:
- No “did the click land?” — Was the intended window in focus? Was the target element at those coordinates? Did the button respond?
- No “did the text appear?” — Was the right input field focused? Did the text render? Did it go to the right app?
- No “did the shortcut work?” — Was the target window elevated (UIPI block)? Did the key combo reach the intended app?
- No “did I actually see what I think I saw?” — OCR returns text, but if a window overlaps between capture and processing, the result is stale by the time it’s returned.
Why it compounds in chain automation:
- A chain step
clickreturnssuccess: trueeven on empty desktop - The next step assumes the prior action worked and compounds the error
- Chain has no built-in “verify after action” step type — poll steps must be authored manually
What partially helps (but doesn’t solve):
find_text_and_click— OCRs first, then clicks. But doesn’t verify the click had an effect.layout_validate— checks stored element positions for drift. Only works on pre-registered layouts.- Chain
pollstep — can wait for text to appear. Requires explicit authoring and knowing what to wait for. uia_invoke— UI Automation has better success reporting. But only works with UIA-compliant apps.wait_for_text— single-tool version of poll.
The core problem:
Verification is left entirely to the AI (call ocr() after click() to check). Most AI agents skip this because they trust the "ok" response, don’t know what to verify against, and can’t afford doubling token cost per action. Chain has no auto-verify primitive.
Planned fix path:
- Add
auto_verifyparameter to action tools (click,type,key_press,type_and_submit,scroll,drag) — when enabled, the tool captures OCR/screenshot context before and after the action and returns a diff/verdict - Return
beforeandafterOCR text + screenshot region in verified action results - Add
verifystep type to chain — wraps any tool call with automatic before/after capture - Add
expectedparameter to action tools — AI specifies what it expects to see after the action (e.g.,click(x=100, y=200, expected_text="Submit"))
Related: L2 — “Tools return ‘ok’ even when the action had no visible effect”
Roadmap / Future Possibilities
R1. Chain interruption — abort mid-sequence
The chain tool runs to completion or global timeout with no stop mechanism. For gameplay chains tied to game state (e.g., react to hit-stun, dodge indicator), the AI needs the ability to abort a running chain and switch to a different sequence.
Approach (not started):
- Add
chain_stoptool that sets an atomic abort flag ChainExecutorchecks abort flag between every step- Poll/loop steps check abort flag before each iteration
- Returns partial result:
{"status": "aborted", "completed_steps": N, "partial_output": {...}}
R2. Database-backed training dataset
AI currently has no long-term memory of what VK sequences worked in previous game sessions. Each session is a cold start.
Approach (not started):
- SQLite schema for gameplay recordings:
recording_id,window_title,game_state_snapshot(OCR text),vk_sequence(JSON),timestamp - Combined OCR + VK logging during keylogger recording
- Queries for replaying sequences that succeeded in similar game states
R3. Custom ML model for adaptive gameplay timing
Done v0.2.46 — Delivered as a Go-native transformer engine (ml/ module). 14K-param transformer trained in-process via Gorgonia. Takes tokenized OCR text + spatial features, predicts optimal tool + coordinates. Persists to model.gob, auto-loads on startup, self-improves from each session’s training pairs. No Python, no ONNX for training.
R4. Smart cropping for OCR performance
Current OCR screenshots capture entire window or full screen. For game UI reading (cooldown numbers, health bars, ability tooltips), this wastes tokens and latency.
Approach (not started):
- Define per-application “UI regions” (e.g., NTE: bottom-center for hotbar, top-left for health)
- Crop screenshot to known UI regions before OCR
- Store region definitions in memory store for reuse across sessions
R5. Video frame analysis
Currently all screen analysis is single-frame (OCR or ONNX on still images). Temporal awareness would enable:
- Detecting game state transitions (loading screen → gameplay, combat → menu)
- Recognizing animation tells (enemy wind-up → dodge window)
- Reading dynamic UI elements (damage numbers, cast bars)
Approach (not started):
- Frame buffer: keep last N screenshots in memory for temporal queries
- Simple state machine: DIFF between consecutive frames to detect large-scale changes
- Video model API integration for frame series → action prediction (long-term)
Notes
- First
UIA FindAllcall costs 16-37s (one-time per process lifetime). Subsequent calls are fast (~280ms children, ~2ms FindFirst). RoInitializenow usesRO_INIT_MULTITHREADED(v0.1.2 fix) to match UIA’sCOINIT_MULTITHREADED— both paths use MTA.- OCR uses native COM WinRT path, falls back to PowerShell on failure.
- Server was built with
-ldflags="-s -w"to reduce binary size.