go-mcp-computer-use

MCP server for Windows desktop computer use

Download as .zip Download as .tar.gz View on GitHub

Known Issues

Safety & Security: Data Collection Controls

See ../security.md for data collection controls, privacy settings, and the full reference including set_config options, watcher management, and noise cleanup.

v0.2.7 — Statistical priors, noise cleanup, config gating

New in v0.2.7:

Feature Status Notes
Statistical prior model (priors_stats) Code complete Element frequency + position per-window, updated on every training save. Go-native, no Python.
Prior confidence adjustment Code complete Applied in ONNXDetect after NMS. Boosts expected elements, suppresses outliers. Gated by prior_adjustment config.
export_yolo_dataset Code complete Exports unused samples as YOLO-format dataset for external training.
training_cleanup_noise Code complete Deletes signal_level=0 samples older than N hours. Supports dry_run.
training_enabled config Code complete When false, disables all auto-save snapshots (actions + watcher). Default: true.
Element priors not yet verified with real accumulated data Untested Priors are updated asynchronously; first detections in a session have no priors loaded until the first loadPriorsFromDB call.
No periodic auto-cleanup of noise Done v0.2.37 retention_days config auto-prunes samples older than N days. Background pruner runs every 6 hours.

Test Session: v0.1.2 — 2026-06-27

Tools Tested & Working (22)

Tool Status Notes
get_screen_size ✅ 3200×900 (virtual — 2×1600)
get_system_info ✅ Hostname, OS, RAM
get_cursor_position ✅ Confirms second screen (x=2142)
get_battery ✅ “no_battery” (desktop)
get_disk_usage ✅ 6 drives enumerated
get_idle_time ✅  
get_keyboard_layout ✅ 00000409 (en-US)
get_network_info ✅ 8 IPs, hostname
get_screen_dpi ✅ 2 monitors, both 96dpi 100%
get_uptime ✅  
get_volume ✅ 49%
list_displays ✅ DISPLAY1 + DISPLAY2 (v0.1.8)
list_windows ✅ 24 windows
list_processes ✅ 195 processes
get_active_window ✅ OpenCode
get_pixel_color ✅ #2b2a33
get_clipboard ✅ “ready”
find_window ✅ OpenCode handle 463694
get_window_state ✅ Visible, maximized, rect (1592,-8,3208,860)
get_display_modes ✅ 37 modes for DISPLAY1
screenshot ✅ Returns base64 PNG
find_text_and_click ✅ Found “OpenCode” and clicked

Tools Tested This Session (additional)

Tool Status Notes
set_clipboard ✅ Write “test from go-mcp-computer-use”
open_url ✅ https://example.com — opens in default browser
scroll ✅ -3 clicks
click (left) ✅ 100,100
click (double) ✅ 200,200 clicks=2
click (right) ✅ 300,300 button=right
click (middle) ✅ 400,400 button=middle
drag ✅ (500,500) → (600,600)
uia_find ✅ GitHub Desktop window found
wait ✅ v0.1.5 — B6: Wait() calc was 1Mx too long
hover ✅ v0.1.5 — B6: same root cause as wait
screenshot_element ✅ v0.1.3 — B5: clamps to screen bounds
uia_get_text ✅ v0.1.7 — B4: nil pattern check added

Tools Not Yet Tested (3)

Tool Reason
type_text Interactive — needs terminal
key_press Interactive
key_sequence Interactive

Test Session: v0.2.27 — 2026-06-30

Tools Tested (v0.2.27 additions)

Tool Status Notes
find_all_images ✅ Degenerate template (1×1 constant color) → NCC skipped → ONNX detected 68+ COCO objects → OCR 200+ words. Valid template at 0.99 threshold → same cascade.
find_image ✅ Degenerate template → ONNX best element returned {x:-19, y:4, w:168, h:47, score:1}.
click (button=middle) ✅ No error on middle click at (100,100).
scroll (horizontal=true) ✅ No error with 3 clicks horizontal scroll.
get_window_state (fullscreen) ✅ Returns fullscreen: false for normal windows, true for borderless fullscreen games.
ocr_languages ✅ Native COM — 2 languages (en-GB, en-US). PowerShell fallback eliminated.

Bug Reports

B1. get_brightness — parse failure (fixed v0.1.4)

Error: parse brightness: strconv.Atoi: parsing "": invalid syntax Root cause: PowerShell/WMI brightness query returns empty string instead of a numeric value. Likely because the display (DVI/HDMI desktop monitor) doesn’t support WMI brightness control (laptops only). Fix: Return a meaningful error or handle gracefully (e.g., "not supported" instead of crash).

B2. list_audio_devices — returns null (fixed v0.1.6)

Result: {"devices":null} Issue: No audio devices enumerated. PowerShell Get-AudioDevice -List may not be installed on this system (requires AudioDeviceCmdlets module). Fix: Return empty slice [] instead of nil slice null.

B3. list_displays — second monitor not enumerated (fixed v0.1.8)

Evidence:

But list_displays only returns DISPLAY1.

Root cause: monitorEnumProc callback gated processing on mi.Flags&1 != 0 (MONITORINFOF_PRIMARY = 0x1). Non-primary monitors were silently skipped. Fix: Removed the primary-only gate — all enumerated monitors are now included, with Primary: mi.Flags&1 != 0 set correctly per-monitor.

B4. uia_get_text / uia_invoke — server disconnect (fixed v0.1.7)

Action:

B5. screenshot_element — negative coordinates rejected (fixed v0.1.3)

Error: x=-8 out of bounds (screen width=1600) Context: Firefox window handle 132490 had rect left=-8 (window decorations positioned off-screen by Windows snap behaviour). Fix: Screenshot element should clamp/clip the region to screen bounds rather than rejecting negative coordinates. A window with off-screen decorations is a normal state (Aero Snap, multi-monitor).

B6. hover — consistently hangs/”Tool execution aborted” (fixed v0.1.5)

Root cause: Wait() used int64(duration) * 10000 (where duration is nanoseconds), producing a timeout 1 million times too long. Wait(100ms) blocked for ~27.7 hours. Fix: Changed to int64(ms) * 10000 (1ms = 10,000 × 100ns intervals). Same fix applies to B7.

B7. wait — “Tool execution aborted” (same as B6, fixed v0.1.5)

B8. find_text_and_click — steals focus

Observation: Calling find_text_and_click brings the target window to foreground. This is expected behavior for a computer-use tool, but worth documenting as a caveat. Workaround: None — by design.

B9. Keyboard input blocked by UIPI on elevated windows (fixed v0.1.9)

Observation: All type, type_and_submit, key_press, select_all_and_type return ok but input does not reach elevated (Administrator) PowerShell. Root cause: Windows UIPI — SendInput with KEYEVENTF_UNICODE from non-elevated MCP server is silently blocked from reaching admin-elevated windows. Fix: Added isForegroundElevated() check using OpenProcess + GetTokenInformation(TokenElevation). Returns clear warning message instead of silent failure.

B10. click may silently fail on elevated windows (documented)

Note: Click uses SetCursorPos + SendInput mouse events. Same UIPI restriction applies — no error feedback when targeting admin windows. Unlike keyboard (which always targets foreground window), mouse targets coordinates, making elevation check impractical. Run MCP server elevated to avoid this.

B11. KeyPress modifier key ordering — Ctrl+C sends C before pressing Ctrl

Observation: KeyPress(["CTRL", "C"]) sends sendUnicode('C') first, then presses Ctrl down, then releases Ctrl. The C arrives before Ctrl is held, so Ctrl+C works as just c in most apps/games. Root cause: KeyPress was splitting keys into three phases: Unicode chars → VK downs → VK ups. Modifiers and their target keys were sent in separate batches, not interleaved. Fix applied: Replaced batch processing with in-order processing. Modifiers are pressed immediately when encountered, then their paired keys are sent while the modifier is held. All pressed modifiers are released in reverse order after the key sequence.

B12. KEYEVENTF_UNICODE may not work in game engines (fixed v0.2.8)

Observation: sendUnicode injects characters via KEYEVENTF_UNICODE, which synthesizes WM_CHAR messages. Game engines using raw input or GetAsyncKeyState for keyboard polling typically don’t see Unicode-injected characters — they check VK codes and scan codes instead. Same issue affects terminals, code editors, and browser input fields in some configurations. Root cause: All keyboard input used KEYEVENTF_UNICODE — letters, digits, TypeText, TypeAndSubmit. Only modifier keys and special keys (arrows, F-keys) used VK codes. Fix: Removed KEYEVENTF_UNICODE entirely. Rewrote keyboard input to use VK codes for everything:


Prompt Engineering: Learn-Once-Reuse-Forever Pattern

The MCP server exposes 161 tools, but an AI agent using them starts cold every session — no knowledge of:

The Pattern

After any successful GUI interaction sequence, the AI should:

  1. Store the sequence as a named macro/recipe (e.g., “open_chrome_search_google”)
  2. Annotate it with application name, window layout details, and screen dimensions
  3. Scope it to the application so it’s reusable across sessions
  4. Next time the same task is asked, recall the stored sequence and execute it directly — no need to rediscover coordinates and timings

Example memory entry

Application: Firefox (v134+ with Multi-Account Containers)
Window size: 1123x791 (positioned at x=295, y=39)
Tab bar: y≈50-70
URL bar: y≈90-110 (click at x=350, y=105 to focus)
Container new-tab: Ctrl+T bypasses popup, or click "No Container" at x=830,y=105
Bookmarks bar: y≈120-140 (when enabled)

Sequence "open_google_and_search":
1. focus_window(handle=132490)
2. click(x=350, y=105) — focus URL bar
3. type_and_submit("google.com")
4. wait(4000)
5. type_and_submit("search query")
6. wait(3000)
7. scroll(clicks=-6)
8. ocr(x=295, y=140, w=1123, h=700) — read results below URL bar

Why this matters

Without this pattern, the AI wastes time and tokens rediscovering basic facts each session — where the URL bar is, that Firefox uses containers, that scroll takes negative values for down. With it, the AI builds a living knowledge base that compounds with every session.

Lessons Learned (from live testing)

L1. Screen layout awareness is critical — always survey before acting

Problem: Commands like click(x,y) / type_text fail silently when the AI doesn’t know the screen layout — what windows exist, their positions, what UI elements are where.

Example from session: Firefox had:

Procedure for any GUI interaction:

  1. get_window_state(handle) — get target window position
  2. ocr(x,y,w,h) over the window region — see what’s actually displayed
  3. Locate the target element (button, text field, link) from OCR coordinates
  4. Click at the element’s center position
  5. Verify with another OCR call after action

Firefox-specific layout (tested v134+):

L2. Tools return “ok” even when the action had no visible effect

type, key_press, type_and_submit, click all return ok — but the input may hit the wrong element or be dropped by UIPI. Always verify with OCR/screenshot after each action.

L3. Firefox containers intercept the + new tab button

Firefox Multi-Account Containers changes the new-tab + behavior — instead of opening a blank tab immediately, it shows a popup asking which container to use. Click “No Container” (≈x=830, y=105 in this layout) or use Ctrl+T which bypasses the popup.

B13. ONNX tools require CGO (disabled in default CGO_ENABLED=0 build)

Observation: go build ./... fails with build constraints exclude all Go files for github.com/yalue/onnxruntime_go when CGO_ENABLED=0. The onnxruntime_go library uses cgo for native shared library bindings.

Impact: ONNX ML tools (onnx_detect, onnx_status, onnx_download) are excluded from CGO-free builds. All other tools work — including the Go-native transformer engine (action prediction), which is pure Go via Gorgonia.

Workaround: Build with CGO_ENABLED=1 and a C compiler:

Status: Documented in v0.2.x. Not a bug — by design. CGO only needed for ONNX; transformer is pure Go.

B14. ONNX YOLO11n model uses unsupported opset 22

Done v0.2.13 — ONNX Runtime DLL updated to v1.26.0 which supports opset 22. Detection works.

B15. No action verification — tools return “ok” for API call success, not for real-world effect

Severity: High — this is the #1 reason the server feels “hit or miss” to an AI agent.

Every action tool (click, type, key_press, scroll, drag, etc.) uses fire-and-forget Win32 APIs (SendInput, SetCursorPos, keybd_event). They return success if the API call didn’t crash — not if the action had any visible effect on screen. There is no built-in verification loop between executing an action and confirming it achieved its goal.

What tools actually return:

The AI’s blind spots:

  1. No “did the click land?” — Was the intended window in focus? Was the target element at those coordinates? Did the button respond?
  2. No “did the text appear?” — Was the right input field focused? Did the text render? Did it go to the right app?
  3. No “did the shortcut work?” — Was the target window elevated (UIPI block)? Did the key combo reach the intended app?
  4. No “did I actually see what I think I saw?” — OCR returns text, but if a window overlaps between capture and processing, the result is stale by the time it’s returned.

Why it compounds in chain automation:

What partially helps (but doesn’t solve):

The core problem: Verification is left entirely to the AI (call ocr() after click() to check). Most AI agents skip this because they trust the "ok" response, don’t know what to verify against, and can’t afford doubling token cost per action. Chain has no auto-verify primitive.

Planned fix path:

  1. Add auto_verify parameter to action tools (click, type, key_press, type_and_submit, scroll, drag) — when enabled, the tool captures OCR/screenshot context before and after the action and returns a diff/verdict
  2. Return before and after OCR text + screenshot region in verified action results
  3. Add verify step type to chain — wraps any tool call with automatic before/after capture
  4. Add expected parameter to action tools — AI specifies what it expects to see after the action (e.g., click(x=100, y=200, expected_text="Submit"))

Related: L2 — “Tools return ‘ok’ even when the action had no visible effect”


Roadmap / Future Possibilities

R1. Chain interruption — abort mid-sequence

The chain tool runs to completion or global timeout with no stop mechanism. For gameplay chains tied to game state (e.g., react to hit-stun, dodge indicator), the AI needs the ability to abort a running chain and switch to a different sequence.

Approach (not started):

R2. Database-backed training dataset

AI currently has no long-term memory of what VK sequences worked in previous game sessions. Each session is a cold start.

Approach (not started):

R3. Custom ML model for adaptive gameplay timing

Done v0.2.46 — Delivered as a Go-native transformer engine (ml/ module). 14K-param transformer trained in-process via Gorgonia. Takes tokenized OCR text + spatial features, predicts optimal tool + coordinates. Persists to model.gob, auto-loads on startup, self-improves from each session’s training pairs. No Python, no ONNX for training.

R4. Smart cropping for OCR performance

Current OCR screenshots capture entire window or full screen. For game UI reading (cooldown numbers, health bars, ability tooltips), this wastes tokens and latency.

Approach (not started):

R5. Video frame analysis

Currently all screen analysis is single-frame (OCR or ONNX on still images). Temporal awareness would enable:

Approach (not started):


Notes