Changelog
[Unreleased]
[0.3.10] - 2026-09-18
Post-0.3.9 computer-use reliability pass: focus/keylogger chain bugs found in live OpenCode testing, plus the ML outcome ledger (desktop-as-teacher). Includes a full local computer-use data reset for clean testing (datalog, training store, memory, transformer weights; ONNX models kept).
Fixed
focus_windowno longer reverse-maximizes windows —FocusWindowcalledShowWindow(SW_RESTORE)unconditionally, which un-maximizes a maximized window every time it is focused (OpenCode observed browsers/editors snap out of maximize). Restore is now applied only whenIsIconic(minimized); otherwiseSW_SHOWkeeps the current maximized/normal placement.- Keylogger/replicate chain replay missing tool —
keylogger_stopemitted"tool":"_focus"andeventsToSmartStepsemittedFocusWindowsteps with an emptyTool. Chain dispatch then failed withunknown tool: _focus/unknown tool:(confirmed inchain_log). Fixes:- keylogger emits
focus_window_by_title - smart steps carry
Tool: focus_window_by_title+focus_windowfield - chain treats focus-only steps (empty Tool + focus field) as successful focus after auto-focus
toolDispatchgainsfocus_window_by_title,_focus(legacy alias),keylogger_start/stop/status,replicate
- keylogger emits
- Local tests no longer require live APPDATA assumptions for ledger schema —
ml_predictionsis created on demand; supervised samples never includeunknown/pending.
Added
ml_predictionsoutcome ledger (desktop as teacher) — predictions fromml_query/agent_suggest(transformer or statistical) are inserted as immutable rows (engine,query/ocr/windowhashes,pred_tool,pred_x/y,confidence,model_version). Only resolution columns update:status,outcome_source,actual_*,verification_data,resolved_at.- Outcome resolver on
click—ResolveMLPredictionsForActionscores pending preds after a real click:- hit — click succeeded within hit radius (~3% screen, min 80px) of a pending coord
- miss — click failed near a pending pred (or same-tool retry nearby)
- unknown — pending rows older than 10m (timeout); never treated as failure
- recovered — a miss followed within ~45s by a successful nearby same-tool action
- Successful clicks far from predictions leave rows pending (agent may have ignored ML) — no poison labels
ml_outcome_statustool +ml_status.outcome— hit/miss/unknown/recovered counts,hit_rate,recovery_rate,unknown_rate,last_hour_hit_rate,eligible_train.- Resolved-only supervised signal —
SupervisedLedgerSamplesreturns onlyhit|miss|recovered; transformer train pool preferstraining_pairs.success=1; unknown/pending never train. - Chain aliases for keylogger/replicate replay —
focus_window_by_title,_focus,keylogger_start/stop/status,replicate(executes provided steps).
Changed
- Clean-slate computer-use data for testing — local
%APPDATA%\go-mcp-computer-usetraining/datalog/memory/ML weights archived then cleared so 0.3.10 testing starts from zero pairs (ONNX models inmodels\retained). Agents must re-collect via normal use;agent_train/scripts/eval-ml.ps1 -Trainrebuild from new data only.
Build / verification
scripts/build.ps1— OK (version 0.3.10, Zig cc + CGO)scripts/lint.ps1—go vetclean + build OKgo test -short ./internal/actions— pass (includes focus/keylogger chain alias tests)
[0.3.9] - 2026-09-18
Fixed
- Transformer training never saw real click coordinates — production
training_pairs.command_jsonstores args as a nested JSON string with mixed-case keys ({"args":"{\"X\":700,\"Y\":400}","tool":"click"}).ml/trainer.decodeCoordsonly unmarshaled top-level lowercasex/y, so on live datalogs every click/hover coord target was(0,0)(measured: 800/800 samples). Fix:ml/dataloader.NormalizeCommandJSON/ApplyNormalizationunwrap string-or-objectargs, acceptX/Y/from_x/to_xcase-insensitively, and fillSample.CoordX/Y+FromCoordX/Y. TrainerdecodeCoordsuses the same unwrap path;makeTargetFromSampleprefers loader-normalized pixels. Unit tests cover the production nested-string shape. - Transformer unusable after restart (empty tokenizer) —
MLEngine.LoadModelcreated a fresh unfitted tokenizer and never loaded vocab.tokenizer.Encodereturnsnilwhen unfitted, soForwardfailed withtoken len 0 != maxLen 128whileIsReady()could still be true; predictions silently fell through to the statistical engine. Fix: training persistsvocab.binnext tomodel.gob(tokenizer Save/Load now wired in production), plusml_meta.json(vocab size, ArgDim/WindowDim/FromCoordDim, tool list, holdout metrics).LoadModelrefuses ready-state without vocab and surfaces a clearlast_error. - Vocab size vs embedding table mismatch — fitted vocabs on real OCR exceeded the hardcoded
VocabSize=2000(observed 2421), so high token IDs were dropped in embedding lookup.modelConfigFor()sizes the model from the fitted tokenizer (rounded up) and is shared by Train / LoadModel / online / finetune paths. - Spatial features were always zero —
prepareBatchcalledencoder.Encode(0, 0)for every sample, and inference used an all-zero coord vector. Training now encodes the sample’s decoded click/drag pixels; inference sets a context point from the live cursor viapredict.Engine.SetContextPoint. - Online training destroyed the tokenizer —
trainFromBufferrefit the vocab on a 32-sample replay batch then ranTrainEpochover the full SQLite DB, scrambling token IDs against existing embeddings. Online updates now train on replay-derived samples withoutFit, keep embedding IDs stable, and re-savevocab.bin+ report real eval loss/accuracy (checkpoint accuracy is no longer hardcoded0.0). - Honest tool-accuracy metric —
trainer.Accuracypreviously took argmax overtoolStart(tools + coord + arg dims), mixing continuous outputs into classification and inflating scores. It now scores argmax over tool logits only, skips unknown tools, and Train logsclick_acc/non_click_accalongside overall holdout accuracy and the majority-class baseline. agent_trainonly rebuilt statistical indexes — it now also trains the Go transformer (writesmodel.gob+vocab.bin+ml_meta.json) and returnsml_statusin the tool payload.scripts/lint.ps1failed when launched fromscripts/—Get-Content VERSIONresolved relative to the script directory. Now uses$PSScriptRoot\..\VERSIONand builds..\cmd\mcp-server.
Added
-
ml_statuson agent ML tools —agent_train,agent_suggest, andchain_predictreturn transformer health:ready,model_loaded,vocab_loaded,vocab_size, paths,last_eval_loss,last_eval_accuracy,majority_baseline,last_error, and asourceverdict (transformerstatistical_preferredunavailable). sourcefield onPredictedAction— predictions are labeledtransformerorstatisticalso agents know which engine produced a coordinate/command guess.- Honest transformer gate —
MLEngine.Predictreturnsnilwhen outputs are near-uniform (no signal) or when last holdout tool accuracy is below the majority-class baseline, soPredictActionsfalls back to the statistical adaptive engine instead of shipping junk coords. cmd/ml-eval+scripts/eval-ml.ps1— local eval harness: tool distribution, coord-decode sanity (with_coords%), majority baseline, optional-Train,ml_statusdump, and smoke predict on real datalog OCR. Exit code2when model accuracy is below majority baseline.- Class balancing for click-heavy logs —
balanceToolsupsamples minority tools (cap 48) before augmentation so a 70%+clickdatalog does not collapse the model to “always click”. Tool one-hot targets are amplified (2.0) so MSE is not drowned by zero coord/arg dims. - Production-shaped ML tests —
ml/dataloader/normalize_test.goandml/trainer/decode_coords_test.golock the nested-args schema that unit tests previously missed (tests used{"x":...}while logs used{"args":"{\"X\":...}"}).
Changed
- README ML section is truthful — the Go-native transformer is described as experimental with vocab/meta persistence, majority-baseline reporting, and an explicit preference for the statistical engine (
ml_query/ml_teach) when the neural path underperforms. OCR/UIA/statistical priors remain the reliable locators. agent_traintool description — documents dual training (adaptive indexes + transformer artifacts) andml_statusin the response.docs/reference/tools.mdregenerated after handler/description updates.- Live retrain on a real datalog (800 pairs) — after the data-path fix,
with_coordswent from 0% → 70.4% for transformer targets;model.gob+vocab.bin+ml_meta.jsonload correctly across process restarts. Holdout tool accuracy on noisy OCR remains at or near the majority baseline — the neural path is gated and labeled rather than silently preferred. Statisticalml_query/agent_suggestnow actually learn token→coord averages from unwrapped production args.
Build / verification
scripts/build.ps1— OK (mcp-server.exe, Zig cc + CGO, version 0.3.9)scripts/lint.ps1—go vetclean + build OKgo test—internal/actionsshort suite +ml/dataloader,ml/trainer,ml/predict,ml/tokenizer,ml/transformerpass
[0.3.8] - 2026-08-27
Fixed
- Transformer training never saw real click coordinates — production
training_pairs.command_jsonstores args as a nested JSON string with mixed-case keys ({"args":"{\"X\":700,\"Y\":400}","tool":"click"}).ml/trainer.decodeCoordsonly unmarshaled top-level lowercasex/y, so on live datalogs every click/hover coord target was(0,0)(measured: 800/800 samples). Fix:ml/dataloader.NormalizeCommandJSON/ApplyNormalizationunwrap string-or-objectargs, acceptX/Y/from_x/to_xcase-insensitively, and fillSample.CoordX/Y+FromCoordX/Y. TrainerdecodeCoordsuses the same unwrap path;makeTargetFromSampleprefers loader-normalized pixels. Unit tests cover the production nested-string shape. - Transformer unusable after restart (empty tokenizer) —
MLEngine.LoadModelcreated a fresh unfitted tokenizer and never loaded vocab.tokenizer.Encodereturnsnilwhen unfitted, soForwardfailed withtoken len 0 != maxLen 128whileIsReady()could still be true; predictions silently fell through to the statistical engine. Fix: training persistsvocab.binnext tomodel.gob(tokenizer Save/Load now wired in production), plusml_meta.json(vocab size, ArgDim/WindowDim/FromCoordDim, tool list, holdout metrics).LoadModelrefuses ready-state without vocab and surfaces a clearlast_error. - Vocab size vs embedding table mismatch — fitted vocabs on real OCR exceeded the hardcoded
VocabSize=2000(observed 2421), so high token IDs were dropped in embedding lookup.modelConfigFor()sizes the model from the fitted tokenizer (rounded up) and is shared by Train / LoadModel / online / finetune paths. - Spatial features were always zero —
prepareBatchcalledencoder.Encode(0, 0)for every sample, and inference used an all-zero coord vector. Training now encodes the sample’s decoded click/drag pixels; inference sets a context point from the live cursor viapredict.Engine.SetContextPoint. - Online training destroyed the tokenizer —
trainFromBufferrefit the vocab on a 32-sample replay batch then ranTrainEpochover the full SQLite DB, scrambling token IDs against existing embeddings. Online updates now train on replay-derived samples withoutFit, keep embedding IDs stable, and re-savevocab.bin+ report real eval loss/accuracy (checkpoint accuracy is no longer hardcoded0.0). - Honest tool-accuracy metric —
trainer.Accuracypreviously took argmax overtoolStart(tools + coord + arg dims), mixing continuous outputs into classification and inflating scores. It now scores argmax over tool logits only, skips unknown tools, and Train logsclick_acc/non_click_accalongside overall holdout accuracy and the majority-class baseline. agent_trainonly rebuilt statistical indexes — it now also trains the Go transformer (writesmodel.gob+vocab.bin+ml_meta.json) and returnsml_statusin the tool payload.scripts/lint.ps1failed when launched fromscripts/—Get-Content VERSIONresolved relative to the script directory. Now uses$PSScriptRoot\..\VERSIONand builds..\cmd\mcp-server.
Added
-
ml_statuson agent ML tools —agent_train,agent_suggest, andchain_predictreturn transformer health:ready,model_loaded,vocab_loaded,vocab_size, paths,last_eval_loss,last_eval_accuracy,majority_baseline,last_error, and asourceverdict (transformerstatistical_preferredunavailable). sourcefield onPredictedAction— predictions are labeledtransformerorstatisticalso agents know which engine produced a coordinate/command guess.- Honest transformer gate —
MLEngine.Predictreturnsnilwhen outputs are near-uniform (no signal) or when last holdout tool accuracy is below the majority-class baseline, soPredictActionsfalls back to the statistical adaptive engine instead of shipping junk coords. cmd/ml-eval+scripts/eval-ml.ps1— local eval harness: tool distribution, coord-decode sanity (with_coords%), majority baseline, optional-Train,ml_statusdump, and smoke predict on real datalog OCR. Exit code2when model accuracy is below majority baseline.- Class balancing for click-heavy logs —
balanceToolsupsamples minority tools (cap 48) before augmentation so a 70%+clickdatalog does not collapse the model to “always click”. Tool one-hot targets are amplified (2.0) so MSE is not drowned by zero coord/arg dims. - Production-shaped ML tests —
ml/dataloader/normalize_test.goandml/trainer/decode_coords_test.golock the nested-args schema that unit tests previously missed (tests used{"x":...}while logs used{"args":"{\"X\":...}"}).
Changed
- README ML section is truthful — the Go-native transformer is described as experimental with vocab/meta persistence, majority-baseline reporting, and an explicit preference for the statistical engine (
ml_query/ml_teach) when the neural path underperforms. OCR/UIA/statistical priors remain the reliable locators. agent_traintool description — documents dual training (adaptive indexes + transformer artifacts) andml_statusin the response.docs/reference/tools.mdregenerated after handler/description updates.- Live retrain on a real datalog (800 pairs) — after the data-path fix,
with_coordswent from 0% → 70.4% for transformer targets;model.gob+vocab.bin+ml_meta.jsonload correctly across process restarts. Holdout tool accuracy on noisy OCR remains at or near the majority baseline — the neural path is gated and labeled rather than silently preferred. Statisticalml_query/agent_suggestnow actually learn token→coord averages from unwrapped production args.
Build / verification
scripts/build.ps1— OK (mcp-server.exe, Zig cc + CGO, version 0.3.9)scripts/lint.ps1—go vetclean + build OKgo test—internal/actionsshort suite +ml/dataloader,ml/trainer,ml/predict,ml/tokenizer,ml/transformerpass
[0.3.8] - 2026-08-27
Fixed
- YOLO UI detection now letterboxes non-square inputs instead of aspect-distorting them -
preprocessYOLOpreviously rescaled X and Y independently to a 640x640 square, so a non-square input (e.g. the 3200x1980 virtual desktop) was squashed by 5x on X and 3.09x on Y, destroying every element’s geometry. The detector then emitted floor-confidence (~0.5) sigma-degenerate 0x0 boxes everywhere. Fix: aspect-preserving letterbox (one uniform scalemin(640/w,640/h), centered gray padding) shared between the encoder and decoder via ayoloLetterboxstruct, andparseYOLOOutputnow inverse-maps boxes with(out - pad) / scale. Verified live: full-screen detection now yields real, varied boxes and click points with confidence 0.5-1.0 instead of thousands of 0x0 floor boxes. elements_query/onnx_detectno longer return garbage off-screen boxes from the letterbox padding -ONNXDetectinverse-mapped detections whose centers sat inside the 640x640 letterbox’s gray padding band back into negative or oversized source coordinates (observed:{948,842,839,236}-> click y:960 on an 860-tall window). Fix: a newclipElementToImagegate filters every detection against the captured bitmap bounds - boxes whose center lands outside the image are dropped (padding false positives), and partially off-screen boxes are clipped to the image soscreen_box/click_pointare always actionable. Verified by unit tests covering all reported garbage shapes.ocr_window/screenshot_elementnow report + key the captured window, not the foreground window -annotateCaptureOptsstampedwindow_titlefromgetActiveWindowTitle()(the FOREGROUND window) regardless of the handle being captured, so handle-scoped captures mislabeled the wrong window and keyed per-window ML-priors from that wrong title. Fix:annotateCaptureOptstakes awindowTitleparam andAnnotateWindowresolvesgetWindowTitle(handle)for the actual captured handle. Verified live:ocr_window(132082)(Mozilla Firefox) now reportswindow_title="Mozilla Firefox"(was “OpenCode”).- Capture responses no longer dump a giant element array on every
ocr/screenshot/onnx_detectcall - the fused YOLO+MobileNet annotation (from 0.3.5/0.3.6) embedded every detected box into every capture response with no cap, so a busy desktop returned thousands of degenerateiconboxes (~1751 observed) and blew the payload to several hundred KB even withinclude_image:false. The AI was forced to grep a monolithic JSON blob. Fix:annotateCaptureOptsnow runsFilterAndCapElementsbefore returning - it drops boxes whose fusedcombined_confidenceis below a floor (default0.40) and caps the surviving list (default50, highest-confidence first). Detection/ML still runs at full fidelity internally; only what’s embedded on the wire is bounded. Verified by unit tests; a 1751-element dump is now at most 50 high-trust rows.
Added
include_imageopt-out for the base64 image bloat in capture-tool responses - everyocr/screenshot/screenshot_element/ocr_window/ocr_active_window/onnx_detect/onnx_classifyresponse embeds the base64 screenshot (~0.7-1.5MB), which wastes model context and forces truncation for text-only AIs. This adds a config flaginclude_image(defaulttrue, preserving current behavior) that globally strips the returnedimage_b64copy while keeping OCR text, element boxes/confidence, and all metadata intact - the internal ML detection pipeline still runs on the frame; only the returned copy is removed. A per-callinclude_imageoverride (true/false) on any capture tool wins over the global config, giving callers a tri-state. Exposed throughset_config/get_configand persisted. Verified live: withinclude_image:falseanocr_windowreturns metadata only (noimage_b64); a per-callinclude_image:truerestores it;onnx_detect(include_image:false)still returns 1677 elements / 350 clickable with the image stripped.exclude_elementsopt-out for text-only captures - a config flagexclude_elements(defaultfalse, preserving current behavior) plus a per-callexclude_elementsoverride onocr,ocr_window,ocr_active_window,screenshot,screenshot_element, andonnx_detectempties the returned element array so a text-only client never drags the detection boxes onto the wire. Likeinclude_image, it flows through the singlesafeHandlerchoke point (stripImageIfExcluded), is persisted viaset_config/get_config, and the per-call override wins over the global value.elements_querytool - a grep-style query over the fused capture for text-based LLMs - instead of dumping the full annotated JSON, the AI asks “where is the submit button” and gets a compact flat projection: filter by source (screen/window/region), YOLOclass, MobileNetlabel, anx,y,w,hregion,min_confidence,clickable_only, and OCRtext(returns matching words with screen coords). Each returned row is just{index, class, label, confidence, clickable, screen_box, image_box, click_point}- the coordinates to click - with no base64 and no nested classifier bloat. Registration mirrors the existingonnx_*tool family.scripts/gen-icons.ps1no longer writes a UTF-8 BOM intowinres/winres.json-Set-Content -Encoding UTF8(PowerShell 7) prepends a BOM, whichgo-winresrejects withinvalid character 'ï' looking for beginning of value, breaking the resource-generation step and hence the whole build. Fix: write with-Encoding utf8NoBOMso the generated JSON parses cleanly.key_presskey names expanded to the full Windows virtual-key surface + case-insensitive lookup + generic modifier combos -vkModMapnow exposes the explicit left/right modifier variants (LCTRL/LCONTROL,RCTRL/RCONTROL,LALT,RALT,LSHIFT/RSHIFT,WIN/META/LWIN/SUPER/CMD/COMMAND,RWIN), andvkSpecialMapgained the numpad block (NUMPAD0-NUMPAD9,NUMPAD_MULTIPLY/ADD/SEPARATOR/SUBTRACT/DECIMAL/DIVIDE), the media/volume keys (VOLUME_MUTE/VOLUME_DOWN/VOLUME_UP,MEDIA_NEXT_TRACK/PREV_TRACK/STOP/PLAY_PAUSE), plusSNAPSHOT,APPS/CONTEXT,CLEAR,EXECUTE, andSLEEP.keyNameToVKnow upper-cases its input so lookups are case-insensitive, and single-char punctuation resolves through thecharToVKtable.KeyPressgeneralized the old CTRL-onlyMOD+keyprefix into any modmap modifier followed by a single letter or digit (CTRL+A,WIN+R,ALT+F4,SHIFT+1, …) and emits the requested modifier’s real VK instead of hardcoding Ctrl. Covered by newinternal/actions/keyboard_test.gounit tests.
[0.3.7] - 2026-08-27
Added
--version,--license, and--helpCLI flags on the server binary -mcp-server.exe --versionprints the build version (go-mcp-computer-use 0.3.7),--licenseprints the full Apache-2.0 text plus theNOTICE, and--helplists the available subcommands/flags. LICENSE and NOTICE are embedded into the binary viago:embed(kept in sync byscripts/build.ps1), so the license text is always retrievable from the shipped executable with no files alongside.
Changed
- Project relicensed from MIT to Apache-2.0 -
LICENSEreplaced with the full Apache-2.0 text, (c) 2026 coff33ninja. Added aNOTICEfile wiring the Apache-2.0 attribution and clarifying that the bundledgpa_gui_detector.onnx(a converted Salesforce GPA-GUI-Detector) remains under the upstream MIT license, not Apache-2.0. README gained a license badge, aLicensesection, and updated model-attribution wording. - Version-info, copyright, and admin manifest now embedded in the executable - resource generation switched from
akavel/rsrc(icon only, no manifest/version) togo-winres(scripts/gen-icons.ps1+ a committedwinres/winres.json.template). The exe now carries aVS_VERSION_INFOblock visible in Windows Explorer Properties -> Details (CompanyNamecoff33ninja, LegalCopyright “Copyright (c) 2026 coff33ninja. Licensed under the Apache License, Version 2.0.”, ProductName, File/Product version fromVERSION) and an application manifest declaringrequireAdministrator(admin is required for the tool’s UIA/UIPI automation to fully work) plus per-monitor-v2 DPI awareness. Previously the exe had no embedded manifest at all (as-invoker, no DPI declaration), so this is an explicit behavioral improvement that enforces the project’s real admin requirement at the OS level. - License decisions: LICENSE + NOTICE are shipped with each release -
.github/workflows/release.ymlnow attachesLICENSEandNOTICEto every versioned release alongsidemcp-server.exe, so the license and third-party model attribution always accompany the binary.
[0.3.6] - 2026-08-27
Added
onnx_classifytool — MobileNetV3 GUI element classifier — a new advisory tier that classifies UI content usingmobilenetv3_small.onnxacross 15 classes (button, checkbox, container, dropdown, icon_button, image, label, link, menu_item, scrollbar, slider, tab, text_input, toggle, unknown). Supportedsourcevalues:screen(full screen),window(active window),region(x,y,w,h),elements(classify every YOLO-detected element crop),crop(raw base64 image). Returns label + confidence for top-N (default 3) per target. Advisory only — it returns an additional signal and never hard-blocks an action.- MobileNet as a chain verification tier — chain
verifysteps accept an optionalclassifyconfig ({enabled, top_n, source}). When enabled, the executed step’s result carries aclassifyadvisory payload with the top classifications, independent of the pass/fail decision. - MobileNet as a watcher advisory tier — the background watcher now classifies each YOLO-detected element crop and attaches
classificationsto its cached detection, giving the AI type+confidence context per element without affecting detection or training. Best-effort and skipped entirely when the model file is absent. AnnotatedCapture/AnnotateScreen/AnnotateRegion/AnnotateWindow— the fused, best-effort annotated-capture production layer. All classifier signals run on the SAME captured frame so geometry stays aligned. Any engine that is unavailable yields an empty sub-block (noted inerrors) without failing the capture.scripts/gpa-gui-export/— redoable GPA-GUI→ONNX conversion — a uv-based project (pyproject.toml+export.py+README.md) that downloadsSalesforce/GPA-GUI-Detectormodel.ptviahuggingface_hub, validates the singleiconclass layout, exports to ONNX(1,3,640,640)→(1,5,8400)opset 12, and writesgpa_gui_detector.onnx.uv sync+uv run export.pyregenerates the artifact from source; it is what CI uses to produce the release asset.- Main README “Models” section — documents the three runtime models, auto-download behavior, and how to regenerate the detector, plus a “UI-aware element detection” feature bullet.
Changed
- Every click now returns an advisory
validatedblock —clickandfind_text_and_clickresults carry a post-click validation combined into MobileNet classification of the click target (top-N label + confidence), the element-priors DB (sample count, prior-adjusted confidence, known-location flag, learned frequency/position), and an ML-memory cross-reference (whether the model has seen this click context before). Hooked at the centralClick()choke point so every click source gets it. Thevalidatedblock is merged at the top level of the tool result so the AI sees a consistent shape whether or not OCR auto-verify ran. Best-effort and non-blocking: if the model or capture fails, the click still succeeds and the validation block simply carries minimal/no data. - Internal refactor —
ClassifyElementsnow delegates to a sharedclassifyElementListhelper (also used by the watcher), keeping crop capture + model inference in one place. - Annotated capture on all AI-facing capture tools —
screenshot,screenshot_element,ocr,ocr_window, andocr_active_windownow always-on return a fusedAnnotatedCapture(OCR text + YOLO element boxes + MobileNet per-element classification + element-priors + ML-memory) with each element exposed in BOTH bitmap-image space and virtual-screen space (screen_box), so text-bound AIs know exactly what is on screen and where to click regardless of whether a vision model is available. Screenshot tools keep the raw b64 as text content for vision AIs while the annotated map rides alongside as structured result.screenshot/screenshot_elementgained an optionallanguageparam passed to the OCR signal. onnx_detectandonnx_classifynow return the same fused annotated capture — instead of separate raw detection/classification payloads,onnx_detectandonnx_classify(screen/window/region) now produce the identicalAnnotatedCaptureshape as every other capture tool, so the classifier signals are aligned on the same frame with dualimage_box/screen_boxcoords,click_point,combined_confidence, and theclickablegate — no redundant double-inference.onnx_detecthonors itsthreshold/iou_thresholdargs;onnx_classifyelements/cropsources return the classification wrapped in the same shape via a shim (no live capture-derived screen boxes in those cases).- The
clickablegate and confidence fusion now trust the MobileNet UI label as the authoritative signal — the YOLO proposal tier emits general COCO object classes (person,vase, …) that are essentially never interactive, so gating on the YOLO class alone always returnedclickable=falseeven when MobileNet had correctly identified a real control (button,text_input,link, …).ElementIsClickablenow accepts an element when either the MobileNet top-1 classified label OR the YOLO class names an interactive type, andCombineElementConfidenceweights an interactive MobileNet label more heavily regardless of YOLO agreement. This makes the fused capture actually recommend clickable controls in the live pipeline. - Coordinate safety + anti-over-click assurance — each annotated element carries a
screen_boxin virtual-screen physical pixels (the exact spaceclickconsumes, no DPI double-scaling), aclick_pointcentroid to press, a fusedcombined_confidence, and aclickablegate (interactive-class whitelist + min confidence) so the AI avoids false-positive clicks on low-trust or non-interactive boxes. The capture also exposeswidth/height,dpi_scale, and the fullvirtual_screenbounds so the AI can reason about screen size and multi-monitor layout. Theml_teach/priors feedback loop weights previously-seen UI higher, so unfamiliar or moved elements are flagged rather than silently clicked. - GUI element detector swapped from generic COCO YOLO to Salesforce GPA-GUI-Detector (single-class
icon) — the previousyolo11n.onnxis a COCO-80 general object detector whose proposals (person,car,bicycle, …) flooded the watcher/priors loop with meaningless, never-interactive regions. The detector is now a UI-native, single-class (icon) ONNX export of Salesforce’s MIT-licensed GPA-GUI-Detector (fine-tuned from OmniParser), which proposes actual interactive-element boxes. The authoritative per-control UI type remains the 15-class MobileNet tier. - Watcher/priors novelty gate now keys on the MobileNet UI label instead of the detector class — because the detector is now single-class (
icon), the watcher dedup and element-priors DB would otherwise bucket every control under one key. The watcher now runs MobileNet classification BEFORE persisting crops and threads each element’s top UI label (button,text_input, …) intoElementKnownConfidently/saveElementRegionSamples, so the priors learn per-control-type locations.DetectedElementgained amobile_net_labelfield carrying the authoritative class. - Detector model auto-downloads when missing —
ONNXDetectnow fetchesgpa_gui_detector.onnxfromhttps://github.com/coff33ninja/go-mcp-computer-use/releases/latest/download/gpa_gui_detector.onnxon first use if absent (best-effort, once per process, serialized against concurrent watcher/tool paths). GitHub redirects that URL to the newest non-draft release’s asset, so it follows every version bump with no hardcoded tag. - Detector model ships with releases —
.github/workflows/release.ymlrebuildsgpa_gui_detector.onnxin CI (via the new export project) and attaches it to each versioned release alongsidemcp-server.exe.
[0.3.5] - 2026-08-27
Changed
- Action-triggered training snapshots are now cropped around the target —
SaveSnapshotAfterAction(used by click, type, drag, hover,find_text_and_click,type_and_submit,select_all_and_type, andclick_menu_item) now captures a 400×400 region centered on the action target instead of saving a full-screen screenshot. This gives the ML model a focused view of the element being acted on, dramatically reducing the size of training data. Older call sites fall back to full-screen capture automatically. - Watcher training samples are now cropped around detected elements — The background watcher (which screenshots the screen every cycle to feed the priors/ML systems) now saves a small padded crop around each detected UI element instead of a full-screen 3200×1980 PNG every cycle. When no elements are detected in a frame, nothing is saved. This prevents the tens of gigabytes of near-identical full-screen images that previously accumulated.
- Watcher is AI-gated with an achievement lock — The background watcher now starts locked (
watcher_locked: trueby default). While locked it still runs ONNX detection and caches results for reference, but never persists training crops. The AI must first prove competence through real interactive actions (click/type/record) and then explicitly callwatcher_unlockto grant the watcher permission to save crops. The unlocked state persists across reboots, so the watcher stays unlocked once earned. - Watcher confidentiality / novelty dedup — The watcher now consults the priors system before saving a crop via a new
ElementKnownConfidentlycheck: an element is considered known when its (class, window) pair has enough samples AND its current normalized location falls within tolerance of the learned position. Familiar elements are skipped; only new/moved/uncertain elements are saved. Snapshots therefore decay as the ML gains confidence, replacing the old fragile count-only gate. - Wallpaper guard — The watcher will never save training crops when the foreground window is the desktop/shell (e.g. “Program Manager”, empty title,
Shell_*, Windows Shell Experience), regardless of the lock state. This prevents the wallpaper from ever being mined as a training signal. - Training DB pruning actually reclaims disk —
PruneOldSamplesnow runsVACUUMafter deleting rows, and a newPruneOrphanedSamplespass (run alongside the periodic retention pruner) removes rows whose image file no longer exists on disk. This reclaims the space that stayed locked insidesamples.dbafter the large full-screen PNG cleanup.
Added
watcher_lock_statustool — reports whether the watcher’s training-crop capture is achievement-locked, plus the count of real-action training samples gathered (action_signal) as guidance on when unlocking is warranted.watcher_unlocktool — grants the watcher permission to persist training crops after the AI has gathered confident data from real interactive actions. The unlock persists across reboots.set_config ... watcher_locked/get_configwatcher_locked— the watcher lock can be inspected and toggled through the standard config surface alongside the dedicated tools.
[0.3.4] - 2026-08-26
Fixed
find_text_and_clickmulti-monitor coordinate bug — OCR returns bitmap-space coordinates (0,0 at top-left of captured image), butClick()expects virtual screen coordinates. On setups with displays above the primary (e.g. Display3 at y=-1080), this caused “y out of bounds” errors. Fixed by tracking the capture region origin (captureX, captureY) and offsetting OCR coordinates before clicking and storing in memory.- Window-scoped OCR coordinate offset —
OCRWindowreturns bitmap-relative coords (0,0 at window top-left), butClick()needs virtual screen coords. Now captures the window rect and usesrect.Left, rect.Topas offset, matching the same fix applied to full-screen and region OCR paths. - Text location memory coordinate consistency —
StoreTextLocationnow stores virtual screen coords (converted from OCR bitmap coords at write time). Memory retrieval uses stored coords directly without re-conversion, preventing double-offset bugs when mixing full-screen OCR and SystemFind sources. - Stale text location memory pruning — On startup, entries with coordinates outside the virtual screen bounds are automatically deleted. Prevents old bitmap-space coordinates (from pre-0.3.4) from being reused. Also added bounds validation in the memory retrieval path — stale entries are skipped even if they haven’t been pruned yet.
Changed
find_text_and_clickscoped window OCR — Whenwindow_titleis provided, now tries window-specific OCR first (viaOCRWindow) before falling back to full-screen. This scopes the search to the target window, avoiding false matches from other windows on screen. Window OCR returns coordinates already in virtual screen space, so no offset conversion is needed.- Watcher auto-starts on boot by default —
watcher_auto_startnow defaults totrue. Previously the watcher had to be manually started withset_configafter every server restart. The watcher feeds the priors system (element position statistics) which improves UI element lookup confidence. Set tofalsein config to disable.
Added
- Z-order layering for text location memory —
TextLocationnow storesz_order(window stack position: 0=topmost).FindTextAndClickcaptures the foreground window’s z-order at call time and uses it for memory matching via newFindTextLocationMatch/FindTextLocationAnyMatchfunctions. These prefer exact z-order matches (clicking the right layer) then fall back to any match. Handles overlapping windows with identical text. - Schema migration:
z_order INTEGER NOT NULL DEFAULT 0column added totext_locationstable. Existing databases auto-migrate viaALTER TABLE. get_configtool — read-only view of all current configuration values plus live watcher status. Returns every field thatset_configcan change, pluswatcher_runningandwatcher_auto_start. Use to inspect state without modifying anything.
[0.3.3] - 2026-08-26
Added
ml_query— ask the ML engine “where is X on this screen?” Pass a query (what you’re looking for) plus current OCR text. Returns coordinate predictions ranked by confidence, matched OCR keywords, and related commands the ML has seen. Searches bothcoordIndex(per-tool coordinate distributions) andwordToCmds(command frequency). Query tokens get priority matching, context tokens add breadth.ml_teach— feed confirmed correct answers back to the ML after every action. Pass what was being looked for, the screen OCR, which tool was used, coordinates, and success/fail. UpdatescoordIndexandwordToCmdsdirectly with both query and context tokens. The learning loop:ml_query→ AI acts →ml_teachreinforces. Each cycle strengthens token→coordinate associations.
Changed
- The ML feedback loop is now complete: query → predict → act → teach. Whether the AI follows an ML prediction or discovers the correct answer itself,
ml_teachensures the ML learns from every outcome — including when the user shows the AI the right answer. - CI workflows updated from
v0.2.xtov0.3.x(ci.yml, auto-tag.yml, jekyll-gh-pages.yml, mod-maintenance.yml). - README status block rewritten — v0.2.x framed as testing/iteration ground, v0.3.x as current stable. Recording & replication and ML feedback loop documented in features section.
docs/architecture.md— addedrecord_replicate.goto code map, ML Loop layer in agent stack diagram, updatedadaptive.goandchain.godescriptions.docs/ci-cd-pipeline.md— updated branch references and diagram to v0.3.x._config.ymllogo URL updated to v0.3.x branch.scripts/gen-tools-doc.go— addedml_query+ml_teachto Adaptive Agent category (now 5 tools).
Fixed
scripts/gen-icons.ps1— rewritten to handle encoding issues and missing dependencies gracefully. No longer fails with PowerShell parse errors whenrsrcis not installed or icon path contains spaces. Build script (scripts/build.ps1) now works end-to-end again.
[0.3.2] - 2026-08-26
Added
- Timed recording no longer blocks MCP —
record(duration_secs=N)now starts a background goroutine for the sleep+auto-stop, returning immediately with a confirmation. Previouslytime.Sleep()inside the handler caused the MCP SDK to kill the request as timed out. Both manual (duration_secs=0) and timed modes now work reliably. - Recording feeds ML on stop —
RecordStop()now callsLogEnrichPatternsFromSession()asynchronously, feeding OCR, UIA, and ML enrichment payloads back to the adaptive engine immediately when recording ends. No manual wiring required — the full loop is: record → stop → enrich → ML learns.
Verified
- 30-second timed recording test: started, returned immediately, auto-stopped after 10 seconds, keylogger inactive — no MCP timeout.
- Manual recording +
record_stop: full session returned with enrichment. OCR snapshots jumped 1668→1817 (+149 from enrichment logging). ML engine received 692 command sequences, 535 click patterns, 26 double-clicks, 26 long-presses, 49 type patterns. Top sequences:"reminder"→click(260 samples),"delete"→click(154),"schedule"→click(152). - AI replication of recorded sessions not yet tested — next validation step.
[0.3.0] - 2026-08-25
Added
- Recording plugin —
record_and_replicate,record,record_stop, andreplicateMCP tools with intelligent chain generation. Recording captures enriched context at each click: OCR text near the click point, UIA element identity (name, automation_id, control_type), ML-predicted coordinates, window title, mouse movements, drags, scroll, and key events. - Typed text capture — keylogger reconstructs typed text from VK code sequences via
vkToCharreverse map. Consecutive printable characters accumulate intotypesteps. Modifier combos (Ctrl+C, Ctrl+V) becomekey_presssteps. Shift+key produces uppercase. Text flushes on focus change, mouse event, or non-char key. - Window snapshot at recording start —
RecordStop()callsListWindows()+GetWindowState()for every visible window.WindowsAtStart []WindowStateInfostored inRecordSession. Replication generatesfind_window→move_window/restore_window/maximize_window+focus_windowusing${var}chain variables to avoid stale handles. - Learning loop —
execToolin chain.go callsLogToolCall(notLogCommand) so every chain step feeds OCR→command pairs back to the ML engine viaLearnFromCommandWithContext.elapsed_mstracked per step for AI pacing analysis. - Smart chain generation —
eventsToSmartStepshandles all event types: type, key_combo (with modifiers), click (all 5 buttons), double_click, long_press, drag, scroll, move_mouse, focus, key_down, key_up. Smart click priority: UIA invoke > OCR find_text_and_click > ML predicted coordinates > raw coordinates. - Context menu detection — right-click generates: click(right) → wait(350ms) →
uia_find(control_type=MenuItem)→find_text_and_click(OCRText)withskip_system_find: true. Menu position cache checked first — if cached item found, skips UIA/OCR and clicks cached item directly. - All mouse buttons — keylogger captures all 5 buttons (left, right, middle, X1, X2). Middle-click →
click(button=middle). X1 →key_press ["Alt","Left"](browser back). X2 →key_press ["Alt","Right"](browser forward). ElapsedMs tracked per click/drag for hold duration. - Double-click detection —
mergeDoubleClicks()post-process merges 2 rapid left-clicks within 400ms at same position (8px tolerance) intoEnrichedEvent.Kind="double_click". Chain generation emits smart click +clickwithclicks: 2. - Long-press detection —
detectLongPress()post-process converts clicks withelapsedMs > 500intoEnrichedEvent.Kind="long_press". Chain generation emits smart click +wait(hold duration in ms). - Menu position memory — global cache:
menuCacheKey{Window, BucketX, BucketY}→[]menuCacheEntry{ItemText, RelativeX, RelativeY, HitCount}. 50px grid bucketing for fuzzy position matching. Hit count increments on repeated observations. Chain tools:cache_menu_items,lookup_menu_items. Thread-safe viasync.RWMutex. - ML integration for enrichment patterns —
double_click,long_press,context_menuregistered as distinct ML tools incoordIndex.LogEnrichPattern()captures OCR context at event coordinates and feedsLearnFromCommandWithContext+RecordResult.LogEnrichPatternsFromSession()batch-logs all enrichment events from a recording session.detectAndLogEnrichPatterns()pre-scans chain steps for enrichment patterns before execution. Wired into bothRecordAndReplicate(after chain execution) andExecuteChain(before chain execution for standalone replicate). uia_invokechain step — chain engine can now execute UIA automation via theuia_invoketool in step sequences.mouse_down/mouse_upchain steps — new chain tools for holding and releasing mouse buttons independently. Used by long-press replay (move→down→wait→up) and available for custom chains.recordMCP tool — starts recording; ifduration_secs > 0auto-stops after that duration, otherwise starts manual mode (userecord_stopto finish).replicateMCP tool — takes a recorded session JSON and replays it as a smart chain with configurableslowdownandloop.record_stopMCP tool — stops an active recording and returns the enriched session with OCR, UIA, and window context at each event.
Changed
- No arbitrary duration limit — AI/user decides when to stop recording.
duration_secs=0enables manual stop mode. - All 5 recording gaps addressed: typed text, window snapshot, learning loop, smart chain generation, context menus.
- All 5 enhancements complete: double-click, long-press, context menu chain, menu position memory, ML enrichment integration.
- Double-click replay uses
clicks: 2parameter with UIA invoke path for reliable replay. - Long-press replay uses
mouse_down→wait→mouse_upsequence for actual button hold duration. - Window snapshot captured at
Record()start so replay restores the layout the user started with. chainUIAInvokereturns error whenUIAInvokereturnsinvoked: false, enabling proper fallback/error reporting.- Shift handling extended —
vkToShiftedCharmap covers digits (1→!,2→@…), OEM keys (-→_,=→+,[→{…), and punctuation. Previously only converted lowercase→uppercase letters. LogToolCallexecutes synchronously to ensure OCR→command bridge state is consistent before the next chain step.enrichClickcoordinates always come from click/drag args — no sentinel values.vkToShiftedCharreverse map built fromcharToVKshift entries in keylogger init.- 178 tests all passing.
Fixed
recordtool null structuredContent —Record(0)in manual mode returned nil session, causing MCP SDK to reject the response. Now returns a confirmation message.
Fixed (pre-0.3.0)
- Nil-pointer in
ensureWindowFocus— addedstate.Rect != nilguard and bounds checking before clicking the window title bar.GetWindowStatecould return a nilRecton certain window types. - CI/CD: tool count auto-patching —
gen-tools-doc.gonow patches the"tools", Nvalue inserver.gostartup log automatically, keeping it in sync with registered tools. - CI/CD: tool categorization — added all 153 tools to
categoryForToolmap with proper categories. Previously 27 tools were uncategorized. - CI/CD: category ordering — added
File OperationsandRecording & ReplicationtocategoryOrderfor consistent doc generation. - Docs: versioning and release checklist — added tool-addition checklist to
docs/reference/versioning-strategy.mddocumenting all files that must be updated when adding tools.
Key Functions
| Function | File:Line | Purpose |
|---|---|---|
vkToChar |
keylogger.go:60 | VK→char reverse map |
vkToShiftedChar |
keylogger.go:65 | VK→shifted char map (!, @, _, +, etc.) |
isModifierVK |
keylogger.go:93 | Check if VK is modifier |
resolveVK |
record_replicate.go:298 | Key name→VK code |
enrichEvents |
record_replicate.go:191 | Raw steps→enriched events + post-processing |
mergeDoubleClicks |
record_replicate.go:420 | Merge 2 rapid clicks → double_click |
detectLongPress |
record_replicate.go:444 | Convert slow clicks → long_press |
cacheMenuItems |
record_replicate.go:33 | Store menu items in position cache |
lookupMenuItems |
record_replicate.go:84 | Retrieve cached menu items |
activeModifiers |
record_replicate.go:135 | Collect held modifier names |
eventsToSmartSteps |
record_replicate.go:569 | Events→chain steps (all types) |
smartClickSteps |
record_replicate.go:558 | Smart click: UIA>OCR>ML>raw |
Record |
record_replicate.go:58 | Start recording |
RecordStop |
record_replicate.go:76 | Stop + enrich + window snapshot |
Replicate |
record_replicate.go:393 | Generate chain with window restoration |
LogToolCall |
datalog.go:208 | OCR→command bridge + ML learning |
execTool |
chain.go:530 | Chain step execution (calls LogToolCall) |
toolDispatch |
chain.go:171 | 53+ chain tools |
LogEnrichPattern |
record_replicate.go:467 | Capture enrichment event → ML + datalog |
LogEnrichPatternsFromSession |
record_replicate.go:507 | Batch-log session enrichment events |
detectAndLogEnrichPatterns |
chain.go:253 | Pre-scan chain steps → ML enrichment logging |
MouseButtonDown |
mouse.go:103 | Send mouse button-down event (for long-press) |
MouseButtonUp |
mouse.go:125 | Send mouse button-up event (for long-press) |
chainMouseDown |
chain.go:1131 | Chain step: mouse_down handler |
chainMouseUp |
chain.go:1143 | Chain step: mouse_up handler |
[0.2.61] - 2026-08-25
Added
record_and_replicateMCP tool — records mouse and keyboard events for a user-specified duration (1-60 seconds), then automatically replays them as a chain. Supportsslowdownfactor (1=real-time, 2=2x slower, etc.),loopcount (repeat the same recording N times), anddelay_ms(pause before replay). Eliminates the manual keylogger_start → wait → keylogger_stop → chain round-trip.
[0.2.60] - 2026-08-25
Fixed
InitAbortFromConfigwired into config handlers — the convenience wrapper forParseHotkeyString+StartAbortPollerwas never called;set_configand startup both inlined the same logic. Now usesInitAbortFromConfigin all 4 call sites, eliminating duplication.makeSequenceTargetswired into training pipeline — the multi-step sequence target builder had zero callers; the trainer only built single-action targets.prepareBatchnow fills the sequence section of the target vector whensequenceLen > 0, enabling the transformer to learn temporal action ordering.
Added
- Sequence-aware training in
prepareBatch— looks aheadsequenceLenconsecutive samples and fills target slots viamakeSequenceTargets. - 6 new tests:
TestInitAbortFromConfig_DisabledNoop,TestInitAbortFromConfig_StartsPoller,TestMakeSequenceTargets_FillsSlots,TestMakeSequenceTargets_ZeroLenSkips,TestPrepareBatch_FillsSequenceTargets,TestPrepareBatch_NoSequenceWhenDisabled.
[0.2.59] - 2026-08-24
Added
system_find_statsMCP tool — returns last-used timestamp and total call count for the system-find feature, enabling AI agents to observe system-find usage patterns.task_is_activeMCP tool — returns whether a task session is currently active (betweentask_beginandtask_end), allowing agents to check session state without callingtask_end.- Text location pruning in retention cycle —
PruneTextLocationsnow runs alongsidePruneOldSamplesin the 6-hour retention pruner, cleaning stale entries from thetext_locationsSQLite table.
Fixed
- Retention pruner goroutine leak on shutdown —
StopRetentionPruner()is now called during server shutdown, closing the background goroutine channel and preventing a goroutine leak on exit.
Tests
- Added 16 tests across 3 new test files:
text_location_test.go(6),system_find_test.go(3),introspection_test.go(3), plus 4 new retention-pruner tests intraining_test.go.
[0.2.58] - 2026-08-23
Fixed
- ML init no longer skipped when adaptive stats training fails —
EnsureAdaptive()returned early from its startup goroutine whenTrainFromDatalog()errored (e.g. transient SQLite lock when many server instances start concurrently), which meantmlEngine.LoadModel()/Train()never ran andchain_predictstayed dead for the entire server lifetime (sync.Onceprevents retry). It now logs a warning and continues, with a 2s backoff before theTrain()fallback. - Augmentation actually reaches training —
Train()calledAugmentAll(samples, 2)but only used the result to fit the tokenizer vocabulary;TrainEpochre-loaded raw samples from SQLite, so augmented pairs never participated in gradient updates. NewTrainer.TrainSamplestrains a caller-provided sample set directly. - Checkpoint accuracy was hardcoded 0.0 —
SaveCheckpointalways received accuracy 0, permanently disablingModelVersion.CheckAndRollback(it requiresbest.Accuracy > 0). Checkpoints now record real holdout accuracy.
Added
- Holdout evaluation for transformer training —
Trainer.Evaluate(mean MSE over a sample set, no weight updates) andTrainer.Accuracy(argmax tool-match against target tool).Train()holds out 10% of datalog pairs before augmentation so the eval set stays honest.
Changed
- Transformer trains multiple epochs on the augmented set — previously one pass over raw samples (loss moved 0.0434 → 0.0436, i.e. not at all). Now 5 epochs over ~3x augmented data with per-epoch eval. Live run on existing datalog (674 pairs): loss_eval 0.0433, accuracy 16.18% (previously unmeasured), saved as checkpoint v3.
[0.2.57] - 2026-08-10
Fixed
- Chain timeout now cancels pending steps —
chainState.ctxis wired to the chain’s cancellation context, andexecStepschecksctx.Done()before dispatching each step. Previously a chain that exceeded its timeout returned failure to the caller while its goroutine kept executing later desktop actions in the background. AddedTestExecuteChainTimeoutStopsLaterStepsproving steps after a timeout never run. - Verified chain steps reuse captured OCR —
execVerifynow callsverifyOCRResults(before, after, cfg)with the OCR snapshots it already took instead of callingVerifyAction, which re-captured full-screen OCR twice more. This eliminates redundant full-screen OCR on every verified step (the cause of a 15s chain timeout during live testing). - Drag uses virtual-desktop coordinates —
Dragnow normalizes absolute mouse input against the virtual screen bounds (viaVirtualScreenBounds()usingSM_XVIRTUALSCREEN/SM_YVIRTUALSCREEN/SM_CXVIRTUALSCREEN/SM_CYVIRTUALSCREEN) and setsMOUSEEVENTF_VIRTUALDESK. Previously it mapped against the primary screen origin(0,0), which broke coordinates on multi-monitor layouts with a negative origin (monitors above or left of the primary). AddedTestNormalizeVirtualDesktopPointUsesVirtualOrigin. - Virtual-desktop-aware bounds checks — coordinate/region validation (
ValidateClickCoord,ValidateRegion), OCR window clamping, andSmartRegionAroundnow use the virtual screen bounds instead of assuming origin(0,0), so tools correctly handle negative-origin multi-monitor desktops.
Changed
CaptureScreennow captures the full virtual desktop (all monitors) rather than the primary display only.VerifyActionrefactored to accept an already-capturedBeforeOCR; newverifyOCRResultshelper shared with the chain verification path.
[0.2.56] - 2026-08-09
Changed
- Dependency security update — bumped
google.golang.org/protobufv1.28.0 → v1.36.11 andgolang.org/x/netv0.46 → v0.56 in both the main andml/modules. Fixes the protojson infinite-loop denial-of-service advisory (CVSS 7.5) in both modules. - Merged Dependabot dependency bumps:
github.com/modelcontextprotocol/go-sdkv1.6.1 → v1.7.0golang.org/x/sysv0.46.0 → v0.47.0modernc.org/sqlitev1.54.0 → v1.56.0github.com/pdfcpu/pdfcpuv0.13.0 → v0.14.0 (main +ml/)
- Closed stale weekly dependency PRs (#8, #22, #44, #48) superseded by the individual bumps above.
[0.2.55] - 2026-07-22
Fixed
-
Transformer broadcasting — biases (
coordProjB,historyProjB,ff1B,ff2B) and layer norm params (ln1W,ln1B,ln2W,ln2B) changed from[MaxLen, d]matrices to[d]vectors, added viaBroadcastAdd/BroadcastHadamardProdinstead of same-shapeAdd/HadamardProd. Softmax normalization replaced ones-matrix outer-product hack withBroadcastHadamardDiv. Reason: the old[MaxLen, d]shapes meant every sequence position learned independent bias/norm parameters — not standard transformer behavior, inflates param count ~24%, breaks generalization across positions, and NAS param-count estimator disagreed with actual model size. -
Multi-monitor wiring —
detectScreenConfig()now enumerates all monitors viaEnumDisplayMonitors+GetDpiForMonitorand populates the spatial encoder’sMonitorsslice. Previously the slice was alwaysnil—MonitorInfostruct andmonitorAt()existed but were never wired up, so multi-monitor DPI was silently ignored. Added 7 tests covering L-shape, vertical stack with negative Y, fallback outside all monitors, and negative-coordinate virtual desktops.
Added
-
Sinusoidal positional encoding — fixed (non-learned) sin/cos positional encoding added to token embeddings before attention.
PE(pos, 2i) = sin(pos / 10000^(2i/d)),PE(pos, 2i+1) = cos(pos / 10000^(2i/d)). Reason: the transformer had zero positional encoding — it was permutation-equivariant over the sequence axis, meaning “click then type” was indistinguishable from “type then click”. Without position information, the model cannot learn temporal ordering of actions. -
Spatial encoder: per-monitor DPI —
MonitorInfostruct with per-monitor rect + DPI scale,monitorAt()lookup, 4th positional pairmonRelX/Y(position within the specific monitor).FeatureDim12→14. Reason: the old encoder sampled DPI at a fixed point(0,0)and applied it globally — on dual-monitor setups with mismatched DPI scales, coordinates on the second monitor got wrong DPI baked in silently.
Changed
- NAS paramCount unified — deleted hand-rolled
paramCount()inml/nas/search.go, now callstransformer.ParamCount(cfg)directly. Reason: the old formula had a phantomMaxLen*d“positional encoding” term (no such params exist) and used3*d*dfor Q/K/V instead of the real model’s4*d*d(Q/K/V/O), so NAS parameter-budget search was scoring against wrong model size.
[0.2.54] - 2026-07-22
Changed
-
Multi-token attention — transformer now processes all
MaxLentoken positions simultaneously instead of averaging embeddings into a single[1,d]vector. Previously Q/K/V operated on one row — attention was triviallysoftmax([1,1]) = 1.0so output was always justV, making Q/K dead weights. Now each of theMaxLentoken positions gets its own embedding, coordinate, and history projection, and attention genuinely mixes information across positions.Attention mechanism (built entirely in the Gorgonia computation graph so gradients flow to Q/K):
scores = Q @ Kᵀ / √d ← [MaxLen, MaxLen] weights = softmax(scores) ← row-wise via Exp → Sum(axis=1) → HadamardDiv output = weights @ V ← [MaxLen, d]Pooling collapses
[MaxLen, d]back to[1, d]for the output head:pooled = Mean(transformer_output, axis=0) → reshape to [1, d] logits = pooled @ headW + headBGorgonia broadcasting constraint: Gorgonia doesn’t support element-wise broadcasting (e.g.
[MaxLen,d] / [d]fails). All biases and layer norms changed from vectors to[MaxLen, d]matrices. Row-wise softmax uses trick:row_sum_2d @ ones[1,MaxLen]produces a[MaxLen, MaxLen]denominator, thenHadamardDivdivides element-wise.Result: Q/K now genuinely participate in attention. Test config: 16K → 27K params. Loss: 0.099 → 0.082 in 20 training steps.
embInputshape[1,d]→[MaxLen,d](per-token embeddings)coordInshape[1,coordDim]→[MaxLen,coordDim](coords repeated per token)coordProjB/historyProjBshape[d]→[MaxLen,d](per-position biases)- Layer norm params
ln1W/ln1B/ln2W/ln2Bshape[1,d]→[MaxLen,d] - FFN biases
ff1B/ff2Bshape[d]→[MaxLen,FFNDim]/[MaxLen,d] - Mean-pool via
gorgonia.Mean(x, 0)+Reshape→[1,d]before output head
- Removed Q/K L2 regularization hack (connected through attention path now)
- Removed Go-side
softmax2D,computeAttentionWeights,getWeighthelpers
[0.2.53] - 2026-07-22
Fixed
- Dashboard URL now prints to stderr on startup so AI can read it.
- Dashboard config check moved inside
Start()— reads config directly instead of relying onActiveConfigtiming.
[0.2.52] - 2026-07-22
Added
- Web dashboard — live monitoring dashboard on random port (auto-picks free port, prints URL on startup). Shows tool usage stats, recent commands, chain history, OCR sequences, training samples, model status. Auto-refreshes every 5 seconds. Configurable via
dashboard_enabledin config. Dark theme, responsive layout.
Fixed
- Model save race condition —
model.Save()andGobSerializer.SaveModel()now use atomic write (write to.tmp, thenos.Rename). Prevents corruption when multiple MCP instances train simultaneously.
[0.2.51] - 2026-07-21
Fixed
training_list_samplesno longer returns massive payloads. List endpoint now returns lightweight metadata only (omitsonnx_detections,normalized_coords,ocr_text,window_rect). AddedTrainingSampleMetastruct.training_cleanup_noisewithmax_age_hours=0now means “delete all noise” (no age filter) instead of silently defaulting to 24 hours.training_cleanup_noisedry-run mode now reportsdeletedcount (files that would be deleted) alongsidefreed_bytes.find_windownow supports case-insensitive substring matching. Tries exact match first viaFindWindowW, falls back toEnumWindowssubstring search. Handles backslashes in window titles.
[0.2.50] - 2026-07-21
Added
- §34 Custom Neural Network (12/12 items done):
- Multi-task output head — model now predicts scroll direction (up/down/left/right) and key category (modifier/navigation/function/alpha/numeric/special) alongside tool and coordinates.
predict.PredictionincludesArgsfield withArgPredictionstruct. Trainer fills arg one-hot targets fromArgsJSON. - Transfer learning —
MLEngine.FinetuneForApp(appName, epochs)fine-tunes the base model on app-specific samples filtered by window title.PredictForApp()uses app-specific model when available, falls back to base.LoadAppModels()restores all app models fromapp_models/directory on startup. - Model compression —
ml/compresspackage with magnitude-based weight pruning (Prune(),IterativePrune()) and weight quantization (Quantize(), 2–8 bit). Compression ratio utilities. Supports prune+quantize stacking. - Neural architecture search —
ml/naspackage with grid search over EmbedDim/NumHeads/NumLayers/FFNDim.Search()evaluates configs,BestConfig()picks best under parameter budget. Parameter count utility.
- Multi-task output head — model now predicts scroll direction (up/down/left/right) and key category (modifier/navigation/function/alpha/numeric/special) alongside tool and coordinates.
- §35 Action Prediction — Tool Expansion, Batch Training, Full Attention:
- Expand tool set —
mlDefaultToolsexpanded from 11 to 25 tools. New:move_window,maximize_window,minimize_window,close_window,find_window,set_clipboard,get_clipboard,type_and_submit,select_all_and_type,ocr,screenshot,click_menu_item,launch_app. OutputDim = tool(25) + from_xy(2) + to_xy(2) + arg(10) + window(6) = 45. - Batch training —
TrainerConfig.BatchSizecontrols mini-batch size. Trainer accumulates samples into batches, callsmodel.ForwardBackward(target)per sample, thenmodel.Step(lr)+model.ResetGradients()once per batch. - Full attention (Q/K/V/O) — Each transformer layer now has separate Q, K, V, O weight matrices (
layerDef{qW, kW, vW, oW}). Scaled dot-product attention computed viaQ @ (Kᵀ @ V / √d)using MatMul-only path to avoid Gorgonia shape squeeze. Layer norm weights changed from[d]vectors to[1,d]matrices for same reason.
- Expand tool set —
- §35 Action Prediction — Sequence Prediction:
- Sequence head —
Config.SequenceLencontrols output: primary = tool + from_xy + to_xy + arg, sequence = N×(tool + to_xy + arg).PrimaryDim(),SeqSlotDim(),TotalOutputDim()layout helpers.predict.SequencePrediction{Primary, Next[]}struct. - Sequence training —
dataloader.LoadSequences(ctx, minLen)groups consecutive training_pairs by session_id.makeSequenceTargets()fills sequence section of target vector in trainer. - Sequence prediction —
PredictSequence()returnsSequencePredictionwith primary + next-N predictions.decodePrimary()anddecodeSlot()methods in predictor. Wired throughMLEngine.PredictSequence()→AdaptiveEngine.PredictSequenceActions(). Helper functions:mlPredictedAction(),predictedCoordFromPrediction(),predictedArgsFromPrediction(). - Chain integration —
chain_predictMCP tool callsPredictSequenceActions(ocr_text), returns primary + future actions as structured JSON for chain execution. - Confidence-gated sequences —
SeqConfidenceThreshold=0.15. Steps below threshold truncate the sequence, providing fallback to shorter prediction when model is uncertain.
- Sequence head —
- §35 Action Prediction — Window Context & Spatial Encoding:
- Window category output —
Config.WindowDim=6, window category one-hot in primary head.WindowCategories= [browser, editor, terminal, file_manager, dialog, other]. Predictor decodesWindowCategory+WindowConffields. Trainer defaults to “other” category. - Window title as input —
PredictWithWindow()/PredictSequenceWithWindow()prepend window title to OCR text before tokenization.chain_predictaccepts optionalwindow_titleparameter. MLEngine + AdaptiveEngine wired. - Spatial relationship encoding — Encoder expanded 7→12 features: added
distFromCenter,isCenter,isEdge,windowAspect,isDialog.spatial.FeatureDim=12.DefaultConfig()updated. - Multi-window awareness —
WindowInfo{Title,Category,Active}struct.PredictWithWindows()/PredictSequenceWithWindows()encode top-3 visible windows as[ACTIVE][category]titleprefix. Wired through MLEngine + AdaptiveEngine. - 4-coordinate output — extended output head from
tool(11)+xy(2)+arg(10)=23totool(11)+from_xy(2)+to_xy(2)+arg(10)=25dims.Config.FromCoordDimfield added. Predictor decodes 4 normalized coordinates. Trainer fills both from/to coordinate targets normalized via spatial encoder. - Action template schema —
decodeCoords()per-tool coordinate extraction: drag uses{from_x,from_y,to_x,to_y}, other tools use{x,y}as destination. Trainer normalizes both pairs through the spatial encoder. - Drag prediction —
Prediction.FromCoordX/Yfields,PredictedAction.FromCoordstruct, bridge populates both coordinate pairs from model output. - Drag-and-drop —
dragtool already inmlDefaultToolswith full from/to coordinate support.drag_and_dropalias handled in trainer’sdecodeCoords(). - Backward compatibility — existing models with old OutputDim (23) fail
LoadParameters(param count mismatch) → gracefully recreated and retrained with new 25-dim layout.
- Window category output —
Changed
ml/transformer/model.go—Config.ArgDim,Config.FromCoordDim,Config.SequenceLenfields.PrimaryDim(),SeqSlotDim(),TotalOutputDim()layout helpers.ml/predict/predictor.go—Prediction.Args,Prediction.FromCoordX/Yfields.SequencePrediction{Primary, Next}struct.PredictSequence()method withdecodePrimary()/decodeSlot().SeqConfidenceThreshold=0.15.ArgCategoriesconstants.ml/trainer/trainer.go—FinetuneEpoch().makeTarget()fills arg one-hot + sequence targets.decodeCoords()per-tool coordinate extraction.Trainer.sequenceLen,Trainer.primaryDim,Trainer.fromCoordDimfields.makeSequenceTargets()fills sequence section.ml/dataloader/loader.go—Sequence,Actiontypes.LoadSequences(ctx, minLen)interface method.ml/dataloader/sqlite.go—LoadSequences()implementation grouping consecutive training_pairs by session_id.internal/actions/ml_bridge.go—MLEngine.PredictSequence()returnsSequencePredictionResult. Helper functions:mlPredictedAction(),predictedCoordFromPrediction(),predictedArgsFromPrediction().internal/actions/adaptive.go—PredictedAction.Args/FromCoordfields.PredictSequenceActions(ocrText)wrapper.PredictActionsWithWindow,PredictSequenceActionsWithWindow,PredictActionsWithWindows,PredictSequenceActionsWithWindows.ml/spatial/encoder.go—FeatureDimexpanded 7→12. New features:distFromCenter,isCenter,isEdge,windowAspect,isDialog.ml/transformer/model.go—Config.WindowDimfield.PrimaryDim()includes WindowDim.ConfigForSize()acceptswindowDim,seqLenparams.DefaultConfig CoordDimupdated to 12.ml/predict/predictor.go—WindowInfo{Title,Category,Active}struct.Predictorinterface extended withPredictWithWindow,PredictSequenceWithWindow,PredictWithWindows,PredictSequenceWithWindows.Prediction.WindowCategory/WindowConffields.WindowCategoriesconstants.encodeWindowContext()helper.ml/trainer/trainer.go—Trainer.windowDimfield.makeTarget()fills window category one-hot (defaults to “other”).internal/actions/ml_bridge.go—WindowInfoForPredictionstruct.MLEngine.PredictWithWindows(),PredictSequenceWithWindows(),PredictWithWindow(),PredictSequenceWithWindow().chain_predicthandler accepts optionalwindow_title.
VERSION
0.2.50 → 0.2.51
[0.2.49] - 2026-07-21
Added
- §34 Custom Neural Network (in progress) — Phase 1–3.3 of the 4-phase ML engine expansion:
- Sequence context —
ml_bridge.Predict()now passes recent action history into the transformer model.AdaptiveEnginemaintains a ring buffer of recent actions, wired throughPredictWithContext(). - Auto-detect screen config — spatial encoder reads actual screen dimensions and DPI via
ScreenSize()+GetDPIScaleForPoint()instead of hardcoded 1920×1080/1.0. - Improved tokenizer — OCR artifact normalization (pipe chars, multi-space),
[NUM]special token for numeric values, character bigram subword fallback for out-of-vocabulary words. Backward-compatible Save/Load. - Online learning —
ReplayBuffer(10K capacity) stores every action execution. Background goroutine retrains the model every 30 seconds from the buffer. - Model versioning —
ml/versioningpackage saves versioned checkpoints (model_vN.gob), auto-rolls back if accuracy regresses >5%, tracks rollback count. Persistence via gob index. - Data augmentation —
ml/dataloader/augment.gogenerates synthetic training samples with DPI scaling (0.75×, 1.25×, 1.5×), OCR noise (common substitutions), and coordinate jitter (±5px). Integrated intoTrain()with 2× augmentation. - Model size presets —
ConfigForSize(small/medium/large)factory for easy architecture experiments.ParamCount()utility. - Debug output —
DebugForward()andDebugInfostruct for model diagnostics (logits, tool probs, norms).
- Sequence context —
Changed
ml/transformer/model.go—Config.HistoryLen,Model.Forward()now accepts history tokens. NewDebugForward(),ParamCount(),ConfigForSize().internal/actions/ml_bridge.go— storesscreenCfg,replayBuf,versioner.RecordExperience()feeds online learning.StartOnlineTraining()/StopOnlineTraining()manage background goroutine. Training uses augmented data.internal/actions/adaptive.go—RecordCommand()now callsmlEngine.RecordExperience().ml/tokenizer/simple.go— bigram subword fallback, OCR normalization,[NUM]token, backward-compatible Save/Load format.
Fixed
- Flaky
TestTrainer_MultipleEpochs— changed strict loss-decrease check to log (allows non-monotonic loss with small datasets).
Removed (this cycle)
- Deleted 6 persona files from repo root (AGENTS.md, SOUL.md, USER.md, IDENTITY.md, TOOLS.md, HEARTBEAT.md).
- Deleted
ml/DESIGN.md— content merged intodocs/reference/models-setup.md. - §33 UPGRADE restored — all 6 tools moved to FAR (nice to have), NEXT emptied.
VERSION
0.2.48 → 0.2.49
[0.2.48] - 2026-07-21
Added
- Smart cascade text finding —
find_text_and_clickandwait_for_textnow use a 3-tier strategy instead of brute-force scrolling:- Spatial memory — checks
text_locationstable for where text was previously seen (position, window, confidence). Instant, no OCR. - System find-text — injects Ctrl+F into the active app (Chrome, Firefox, Edge, Brave, Opera, Notepad, VS Code, Explorer) and OCRs the match highlight. ~500ms, zero scrolls.
- OCR + scroll — falls back to screenshot OCR with auto-scrolling only if the first two tiers miss.
- Spatial memory — checks
text_location_memory.go— SQLite-backed spatial RAG store. Every successfulfind_text_and_clickauto-saves word/line text, screen position, window title, and hit count. FTS5 index for fast lookup. Confidence increases with repeated hits.system_find.go— OS-level find-text viaSendInputCtrl+F injection. Tracks usage stats (last used timestamp, total count). 5-second timeout guard on the full Ctrl+F workflow.find_text_smart.go— unifiedSmartFindText()andSmartFindTextAndClick()orchestrating the memory → system-find → OCR cascade. New params onfind_text_and_click:window_title(for system-find),skip_memory,skip_system_find.WaitForTextScroll—wait_for_textnow scrolls when text isn’t visible (passmax_scrolls> 0), using the same cascade.
VERSION
0.2.47 → 0.2.48
[0.2.47] - 2026-07-21
Fixed
- Go 1.26 panic on startup — updated
go4.org/unsafe/assume-no-moving-gcto latest version. Exe no longer requiresASSUME_NO_MOVING_GC_UNSAFE_RISK_IT_WITHenv var.
VERSION
0.2.46 → 0.2.47
[0.2.46] - 2026-07-21
Added
- Go-native ML prediction engine — real transformer neural network (Gorgonia) replaces statistical word-counting for UI automation predictions. 9-package
ml/sub-module: transformer (FFN + residuals + Adam optimizer), tokenizer, spatial encoder (7-feature DPI-aware), SQLite dataloader, training pipeline, softmax/sigmoid inference. Model trains fromtraining_pairsSQLite table on first run, persists tomodel.gob, loads on restart. Falls back to existing statistical engine when ML model is unavailable. ml_bridge.go— integration layer wiring ML transformer intoAdaptiveEngine.PredictActions()tries ML inference first, falls back to statistical engine.EnsureAdaptive()loads or trains model on startup.- Self-improving per-machine models — each installation builds its own personalized model from local usage history. No pre-trained model shipped.
launch.ps1— launcher wrapper that setsASSUME_NO_MOVING_GC_UNSAFE_RISK_IT_WITH=go1.26automatically. Use this or add the env var to your MCP client config’senvfield (all examples updated).
VERSION
0.2.45 → 0.2.46
[0.2.45] - 2026-07-19
Fixed
- Nullable union type serialization bug — the Go jsonschema library generates
"type": ["null", "integer"]for*int32fields and"type": ["null", "boolean"]for*boolfields. The opencode MCP client cannot serialize values for nullable union types, producing truncated JSON (e.g.{"x":instead of{"x": 100}). This brokescreenshot,ocr,find_text_and_click,image_diff, and any tool withVerifyArgs(auto_verify,expected,pre_expected). Fixed viaaddToolCleanwrapper that auto-generates clean schemas with plain"type": "integer"(not nullable) for all 100+ tool registrations. Recursively strips nullable unions from nested properties, array items, definitions, and additional properties.
VERSION
0.2.44 → 0.2.45
[0.2.44] - 2026-07-19
Added
image_difftool — pixel-level screenshot comparison. Accepts two base64 PNGs (before,after) and returnschanged_pixels,total_pixels,change_ratio(0-1),mean_diff(0-255),max_diff(0-255),same(bool). Optionalthreshold(0-255, default 30) controls per-channel sensitivity. Optionalgenerate_imagereturns a diff image with changed pixels highlighted in red. Handles mismatched dimensions by comparing the overlapping region. Lives inverify.goalongsidecomputeTextDiff.
VERSION
0.2.43 → 0.2.44
[0.2.43] - 2026-07-19
Added
- File-based structured logging — rotating JSON log files at
%APPDATA%/go-mcp-computer-use/logs/server.log. Configurable viaset_config:log_file_enabled(default: true),log_file_max_size_mb(default: 10),log_file_retention(default: 7 rotated files). Logs persist across restarts. get_logstool — reads server log entries from the file-based log. Supports filtering bylevel(min level),search(keyword),since_minutes(time window), andlines(max entries, default 50, max 500). Returns parsed log entries with timestamps, levels, messages, and attributes.report_issuetool — generates a GitHub issue report with system info, recent error logs, and AI-provided context. IfghCLI is available, creates the issue automatically with thebuglabel. Otherwise returns the full markdown body for manual submission.- Panic recovery wrapper —
safeHandlergeneric wrapper catches panics in tool handlers, logs the panic value + full stack trace, and returns a structured MCP error instead of crashing the process. Applied toget_logsandreport_issue. Global panic recovery inmain.gocatches any remaining panics. - Multi-handler slog — logs now go to both stderr (TextHandler, for MCP transport) and file (JSONHandler, for persistence) simultaneously. File handler uses the configured
log_level. - New config fields —
log_file_enabled(default: true),log_file_max_size_mb(default: 10),log_file_retention(default: 7). All persist viaset_config.
VERSION
0.2.42 → 0.2.43
[0.2.42] - 2026-07-19
Added
chain_aborttool — checks if the global chain abort hotkey has been pressed since last check. Returns{aborted: true}when the configured key combo (default: Ctrl+Shift+Escape) is detected. The abort is consumed on read (auto-resets). Call before starting long chains or poll periodically.set_window_locktool — locks the active chain to a specific window by handle. Screen-touching tools (click, type, OCR, etc.) will verify the locked window is foreground before executing. Ifwindow_lock_auto_focusis enabled, automatically re-focuses the locked window when it loses foreground.clear_window_locktool — releases the window lock. Screen-touching tools will no longer be restricted to a specific window.- Chain abort hotkey engine — background goroutine polls
GetAsyncKeyStateevery 50ms (configurable) for a global hotkey combo. When all keys in the combo are held simultaneously, the abort channel closes and any running chain terminates at the next step boundary. - Window lock-on context —
verify_window_lock()is called before every screen-touching step in chains whenwindow_lock_enabled=true. Prevents accidental actions on wrong windows when focus is stolen. - New config fields —
chain_abort_enabled(default: true),chain_abort_keys(default: “Ctrl+Shift+Escape”),chain_abort_poll_ms(default: 50),window_lock_enabled(default: false),window_lock_auto_focus(default: true). All persist viaset_config. go-mcp-computer-useskill for Claude — new skill file that teaches Claude about this MCP server’s tools, architecture, and limitations.scripts/test.ps1— PowerShell test runner with Zig CGO setup (mirrorsscripts/build.ps1),-Integrationflag,-Shortflag, and configurable timeout.SaveToBytes/LoadFromBytesconfig helpers — enables testing config round-trips without touching the filesystem.
Fixed
- Chain goroutine leak on global timeout —
ExecuteChainspawned a goroutine forexecStepsbut never cancelled it on timeout. The goroutine would run to completion (or hang) after the parent returned. Now usescontext.WithCancel— timeout cancels the context, and the abort case also cancels. - Chain abort channel shared across chains — previously the abort channel was never reset after being consumed, so a single hotkey press would abort all future chains. Now
GetAbortChanneldetects a closed channel and creates a fresh one automatically.
VERSION
0.2.41 → 0.2.42
[0.2.41] - 2026-07-18
Added
reset_statetool — clears accumulated adaptive engine stats (timings, successes, sequences, word→cmd index, coord index) and OCR→command bridge buffer. Use between heavy batch operations to prevent state accumulation and MCP timeouts.dismiss_all_menustool — presses Escape to dismiss open context menus/dialogs. OCRs before and after to detect which menus were open and whether they successfully closed. Returns{esc_pressed, menus_before, menus_still_open}.- Error enrichment for
find_text_and_click— when target text is not found, the error now includes up to 15 visible lines from OCR so the AI can see what IS on screen (e.g."text 'Delete...' not found. Visible text: 'Back | Refresh | Save as | Print...'").
Removed
AI_ERROR_REPORT.md— error report findings addressed in v0.2.41 (reset_state, error enrichment, dismiss_all_menus).
VERSION
0.2.40 → 0.2.41
[0.2.40] - 2026-07-15
Added
cmd/credit-audit— MCP tool credit audit — new Go tool that measures JSON payload size and estimated token cost of every read-only MCP tool. Covers 66 probes across 14 groups (System, Display, Audio, Window, File, Perception, ONNX, UIA, Memory, Template, Training, Datalog, Agent, Layout). Features per-probe 60s goroutine timeout to survive hanging probes,-jsonflag for machine-readable output, and-sessions Nto project hog-tool costs across N calls. Identifies the top credit hogs:training_list_samples(4.2 MB — base64 screenshots in training DB),record_screen(1.1 MB),screenshot_full(580 kB),priors_stats(121 kB).
Fixed
loadPriorsFromDBreentrant mutex deadlock —UpdatePriorsFromDetectionsandUpdatePriorsForNegativeheldelementPriors.mu.Lock()then calledloadPriorsFromDB()which tried toLock()the same non-reentrant mutex. Split intoloadPriorsFromDB()(acquires lock) andloadPriorsFromDBLocked()(assumes caller holds lock). The two write-path callers now use the locked variant directly.- Fragile
RLock→RUnlock→Lock→RLocklazy-init pattern —AdjustConfidenceWithPriors,GetPriorStats, andFindPriorPredictioneach manually dropped the read lock, acquired the write lock, then re-acquired the read lock. Replaced withsync.Oncefor the initial load, eliminating the lock-dance and preventing future reentrant deadlocks if anyone adds a code path between the drop and re-acquire. TaskBeginnested lock ordering — heldtaskMu.Lock()while acquiringdlogMu.Lock(), creating ataskMu→dlogMuordering dependency. Restructured to releasetaskMubefore acquiringdlogMu, matchingTaskEnd’s pattern.Analyzehelde.mu.RLockacross external lock acquisitions —Analyze()calledtt.Stat()andts.Rate()while holdinge.mu.RLock(). IfStat()orRate()ever acquirede.mu, this would deadlock. Fix: snapshot all maps under RLock, deep-copyPersistedStatvalues (not just pointers), release RLock, then compute stats.LogToolCallimplicitpairMudependency in log parameter —bridgeBufferSize()(which acquirespairMu) was called as an argument toslog.Warn, evaluated in the parameter list before the explicitpairMu.Lock()later in the function. Extracted to a variable before the log call to decouple lock acquisition ordering.
VERSION
0.2.39 → 0.2.40
Added
get_dpi_for_point(x, y)— new MCP tool that returns DPI and scale percentage at a specific screen coordinate. Useful for determining which monitor a coordinate is on and its scaling factor in mixed-DPI multi-monitor setups. Returns{dpi, scale_percent, x, y}. Also available in chain dispatch.FocusHandlechain step field — newfocus_handlefield on chain steps accepts a handle directly (via `` capture), avoiding title re-resolution. Chain checksfocus_handlefirst, falls back tofocus_window(title-based).auto_verify_focuschain option — new boolean field on chain requests. Whentrue, the chain engine tracks the last-focused window handle and re-verifies foreground state before sending input (click, type, key_press, scroll, etc.). If focus was stolen by a popup or notification, it re-focuses before acting.click_menu_itemaccepts handle — new optionalhandlefield. When provided, bypasses title-based window lookup. Falls back towindow_titleif handle is 0.layout_validateaccepts window handle — new optionalwindow_handlefield. When provided, bypasses title-based window lookup. Falls back towindow_titleif handle is 0.
Fixed
FocusWindownow verifies and retries — previously discardedSetForegroundWindowreturn and never confirmed the window actually became foreground. Now uses a 4-attempt fallback chain: (1)SetForegroundWindowwithAttachThreadInput, (2)BringWindowToTop+SW_SHOW, (3)SwitchToThisWindow(bypasses foreground lock), (4) retrySetForegroundWindowafter delay. Each attempt verified viaGetForegroundWindow. Returns error if all attempts fail instead of silently returning nil. This fixes the stale-focus failure where clicks land on the wrong window.click_menu_itemno longer silently fails — when window title matches but window is obscured, the improvedFocusWindow(called by browser focus helpers) now actually brings the window to foreground before OCR+click.ensureWindowFocustitle-bar click guarded by z_order — the post-focus activation click atTop+10was blind to always-on-top windows. Now checksZOrder == 0(truly topmost) before clicking, preventing accidental hits on overlapping windows.
VERSION
0.2.38 → 0.2.39
[0.2.38] - 2026-07-15
Added
z_orderfield onget_window_state— reports 0=topmost, higher=deeper in Z-order stack. UsesGetDesktopWindow+GW_CHILDto find the true topmost, then walksGW_HWNDNEXTcounting visible windows. The AI can compare z_order between handles to determine absolute stacking: window A at z_order=3 is above window B at z_order=12.uia_get_all_elements(handle, max_results)— returns all immediate child UI elements in a window (title bar, menu bar, content panes, toolbars, status bar). UsesTreeScope_Children+TrueCondition(one level deep, not recursive DOM tree) so a browser window doesn’t flood with 10K+ elements. Returns name, control_type, automation_id, bounding rect, is_enabled for each.uia_get_element_at_point(x, y)— identifies which UI element is at screen coordinates using UIAElementFromPoint. Returns name, control_type, automation_id, bounding rect. Use afterget_cursor_positionor click to validate what was under the cursor.wait_for_ui_element(handle, name, control_type, timeout_ms)— polls UIA FindFirst on a window’s descendants until the element appears or timeout. Use for content verification: click a button, then wait_for_ui_element for the dialog that should appear. Default timeout 10s.- Auto-capture element_at_point in chain — mouse-based chain steps (
click,move_mouse,hover,drag) now automatically callUIAElementFromPointat their target coordinates after execution. The UIA element is attached to the step output aselement_at_pointand available as a captured variable for subsequent steps. verify_uistep type — new chain step that verifies a UI element exists or disappears using UIA instead of OCR. Acceptselement_name,control_type,handle(window scope),timeout_ms,not_exists(true = expect absence). Polls UIAFindFirst/WaitForUIElementuntil timeout. Complements the existing OCR-basedverifystep for structural post-action validation.if_uiastep type — new conditional branch that checks UI element existence via UIA. Branchesthen/elsebased on whether an element with the given name/control_type is found. Likeif(OCR), but structural rather than pixel/text-based.- Chain-callable UIA tools —
uia_find,uia_get_element_at_point,uia_get_all_elements,uia_set_text,wait_for_ui_elementare now registered in the chain tool dispatch and can be called as regular chain steps with `` substitution. - Chain-callable window tools —
get_active_window,ocr_window,ocr_active_windowadded to chain dispatch. ocr_windowtool — new MCP tool that extracts text from a specific window by handle using Windows OCR. CallsOCRWindow(hwnd, language)which captures the window’s bounding rect viaGetWindowRectByHandle, clamps to screen bounds, then runs WinRT OCR on the region. Acceptshandle(uintptr) and optionallanguage.ocr_active_windowtool — new MCP tool that extracts text from the current foreground window. UsesForegroundWindowHandle()to get the active window handle, then delegates toOCRWindow. Accepts optionallanguageparameter.ForegroundWindowHandle()exported helper (datalog.go) — returns the current foreground window’s HWND viaGetForegroundWindow.clamp32()helper (ocr.go) — clamps int32 values to [lo, hi] range for safe screen bounds capping.list_windowsnow returns bounding rect — each window now includesx,y,width,heightfromGetWindowRectByHandle. The AI can cross-reference window bounds withlist_displaysmonitor positions to determine which screen a window is on.get_active_windownow returns bounding rect — samex,y,width,heightfields added.
Changed
OCRWindownow clamps to screen bounds — window rects with negative coordinates (off-screen title bars) are clamped to the virtual screen dimensions viaScreenSize()before capture, preventing “out of bounds” errors fromValidateRegion.LogToolCallauto-capture unchanged — remains full-screenOCRScreento preserve the training pair bridge behavior; the AI chooses betweenocr,ocr_window,ocr_active_windowas needed.- Tool descriptions updated —
get_window_statedocuments z_order and the visible-flag vs actual-visibility distinction.focus_windowtells AI to ALWAYS call before interacting with a non-foreground window.uia_findmentions it can find textboxes, address bars, search menus, title bars; advises focusing window first.ocr_windowwarns about minimized/obscured windows.list_windowsandget_active_windowdescriptions mention bounding rect fields. - Chain input schema relaxed —
argsAdditionalPropertiesnow accepts any type ({}instead of{type:object}), so `` template strings in arg values pass schema validation.
VERSION
0.2.37 → 0.2.38
[0.2.37] - 2026-07-08
Added
- Per-tool enable/disable —
tool_denylistconfig field (array of tool names) removes tools from the MCP server entirely so the AI never sees them. Usesserver.RemoveTools()after registration — zero changes to existing handler code. Case-insensitive matching. Example:"tool_denylist": ["shutdown", "restart", "hibernate"]. Configurable at runtime viaset_config. - Retention policy —
retention_daysconfig field (integer, 0=disabled) auto-prunes training samples older than N days. Background pruner runs every 6 hours. Deletes both database rows and image files. Starts automatically on boot whenretention_days > 0andtraining_enabledis true. Configurable at runtime viaset_config. - Unit tests for v0.2.33 behaviors — 25 new tests across
internal/actions/adaptive_test.goandinternal/actions/datalog_test.go:uniqueTokensdedup,nearbyOCRTextspatial scoping,capAndDedupeTextfallback,SaveAdaptiveStat/LoadPersistedStatsround-trip,Analyze()persisted fallback. 5 new tests ininternal/config/config_test.goforToolEnabledhelper. ToolEnabledhelper (config.go) — case-insensitive denylist check with empty-string safety.
Changed
set_configtool — accepts new fields:tool_denylist(string array),retention_days(integer). Description updated to document both.Configstruct — addedToolDenylist []stringandRetentionDays intfields.
[0.2.36] - 2026-07-07
Fixed
- All handlers now return structured JSON instead of hardcoded
"ok"—verifiedResult()helper and ~50+ handler functions setContent: []mcp.Content{&mcp.TextContent{Text: "ok"}}which blocked the MCP SDK from auto-populating with the structured second return value. WhenContent == nil, the SDK marshals the second value and sets bothStructuredContentandContent(as JSON text). FixedverifiedResult()to returnnilContent when extra data is present, and changed every handler that returned explicit"ok"Content to return&mcp.CallToolResult{}, map[string]any{"ok": true}, nil(no-data tools) or&mcp.CallToolResult{}, result, nil(data tools). Verified live post-reboot:list_windows,get_system_info,get_active_window,ocrall return their structured JSON payloads instead of"ok".
[0.2.35] - 2026-07-07
Fixed
roUninitializenever called on shutdown — COM WinRT apartment initialized byensureRo()was never cleaned up. AddedCloseWinRT()inocr_com.gowired intomain.go’s signal handler soRoUninitializeruns on graceful shutdown.statusinitializer unreachable inverifiedResult—status := "ok"was always overwritten by the if/else branches before use. Changed to single defaultstatus := "ok (verified)"with negation on!vr.Passed.klMouseDownandwinEventHookProcdead code — leftover scaffolding from keylogger refactor.winEventHookProctype was unused (callback usessyscall.NewCallbackinline);klMouseDownwas shadowed by local vars inpollLoop. Removed.- Unused
stateparams inexecVerifyandexecPoll— both functions received*chainStatebut never used it. Renamed to_ *chainStateto match the pattern already used byexecWaitandexecTool. Keeps call-site interface compatibility withexecIf/execLoopwhich do need state. - Inefficient string concatenation in
WriteStringcalls — three spots inreadXlsxandreadPdfconcatenated strings before passing tobuf.WriteString(...), allocating intermediate strings. Changed tofmt.Fprintf(&buf, ...)which avoids the extra allocation.
Added
uia_set_textMCP tool — writes text into a UI element via UI Automation’sValuePattern.SetValue. The COM plumbing (uiaElement.setValue) existed but was never wired. AddedUIASetText()inuia.go, handler and registration inserver.go. Fills a backlog item fromdocs/meta/backlog.md.- Prior-based coordinate prediction in
FindUIElement—priors.gotracked element frequency and position per window butFindUIElementnever consulted them. AddedFindPriorPrediction()that returns predicted coordinates for high-confidence priors (frequency >= 70%, sample_count >= 5, StdX/StdY <= 2.0). Inserted as step 1.5 in the cascade: memory → prior → ONNX → OCR. A prior hit avoids ONNX (no GPU/CPU inference) and OCR (no PowerShell launch). - ONNX/OCR failure logging in
FindUIElement— when ONNX detection failed, the function silently fell through to OCR with no trace of the error. Same for OCR failure. Addedlog.Printfcalls in both paths so server logs record when these subsystems are unavailable, distinguishing “not found” from “can’t detect.” - Competitive intelligence gathered from landscape survey — cross-referenced 12 open-source and commercial projects in
docs/comparison-vs-alternatives.mdand extracted features worth adopting. Added a “Competitive Intelligence” section todocs/meta/plan.mdwith prioritized steals (quick wins → medium → high effort). Fleshed outdocs/meta/backlog.mdwith 3 new sections (Memory & ML, Transport & Server, Browser Automation, Linux & Container) and enhanced Security & Identity, Vision, Mouse, Screen, and Processes sections — totaling ~41 new backlog items. Shout out to Cua, Agent-S, Bytebot, MS Magentic-UI, Windows-MCP, DesktopCtl, Windows MCP Server, Computer Control MCP, Browser Use, and Microsoft Windows Recall for the reference implementations and design inspiration.
Fixed
- Double screenshot in
FindUIElementOCR fallback — when ONNX failed,OCRScreen("")captured the screen again even thoughCaptureScreen()had already been called at step 2. Changed OCR fallback to callocrFromBase64(b64, "")directly, reusing the existing screenshot. Also performspushRecentOCR/tryCompletePair/LogOCRSnapshotside effects thatOCRScreennormally handles. - Fragile memory deserialization in
FindUIElement— the memory fast path used 4 levels of nested type assertions (any → map → float64 → int32) with silent fall-through on any mismatch. Extracted intomemoryToElement()helper that returnsnilon any parse failure, keeping it clear and testable.
[0.2.34] - 2026-07-07
Fixed
- Training pair pipeline died permanently after a single missed OCR window —
LogToolCallonly calledOCRScreen("")(the auto-capture that completes a pending pair and refreshesrecentOCRfor the next action) inside theif ocrBefore != ""branch. If one action’sfindRecentOCRBeforecall missed thebridgeWindow— entirely plausible under normal agent round-trip latency between an explicitocr()call and the action that follows it — the buffer never refreshed again. Every subsequent action would also find nothing, permanently starvingtraining_pairsuntil something called an OCR tool manually to re-seed it. Moved theOCRScreen("")auto-capture out of the conditional so it now runs after every action regardless of whether that action found a prior snapshot, making the bridge self-healing instead of a single point of failure.
Changed
bridgeWindowwidened 30s → 60s — 30s was tight for realistic agent-driven latency between an OCR call and the action that follows it. Combined with the self-healing refresh above, 60s gives headroom without letting stale context linger indefinitely.
Verification
Live end-to-end test through the MCP server (ocr → click → type → key_press → click → type): before this fix, 5 real actions produced 0 training_pairs rows (the bridge had already gone dark from an earlier missed window). After the fix and a rebuild/restart, the same kind of sequence produced 5/5 pairs, with agent_train/agent_analyze showing populated timing_stats, success_rates, and top_sequences with counts correctly bounded by total_commands.
[0.2.33] - 2026-07-06
Fixed
-
agent_analyze: timing_stats/success_rates resetting on restart — Stats were purely in-memory (map[string]*ToolTiming,map[string]*ToolSuccess), so every server restart wiped them, leavingagent_analyzewith empty timing_stats/success_rates until new commands arrived. Addedadaptive_statsSQLite table andSaveAdaptiveStat()/LoadPersistedStats()/HydratePersisted(). On startup,EnsureAdaptive()callsHydratePersisted()to seed in-memory stats from the durable table, and everyRecordResult()call persists the aggregate asynchronously viago SaveAdaptiveStat(...). -
Training token inflation:
uniqueTokens()dedup —tokenize()on raw OCR text produced duplicate word entries per row (same word appearing multiple times in the OCR dump contributed multiple hits to the (word, command) pair). This causedCountintop_sequencesto exceedtotal_commands/total_sequences(which count rows, not word occurrences). AddeduniqueTokens()to dedupe before insert, so each (word, command) pair gets at most 1 hit per row regardless of how many times the word appears in the OCR context. - OCR context scoped to nearby words —
findRecentOCRBefore()used to pass the ENTIRE screen’s OCR dump asocr_beforefor every command, so common on-screen words (“the”, bullet points, news headlines, unrelated UI labels) all got associated with whatever command followed — producing noise-dominated predictions. Now:- For coordinate-based tools (click, move_mouse, drag, hover), it uses word bounding boxes (
words []OCRWord) to scope context to words within 200px of the target coordinates, nearest-first, capped at 20 words. - For non-coordinate tools, it falls back to
capAndDedupeText()which returns at most 40 unique words from the full dump instead of the raw dump.
- For coordinate-based tools (click, move_mouse, drag, hover), it uses word bounding boxes (
LogToolCallno longer double-counts timing/success —LogToolCallpreviously calledAdaptive.RecordResult(tool, 0, errVal == nil)with a phantom 0ms duration. Callers (Click, TypeText, etc.) already callAdaptive.RecordResultwith the real elapsed duration. Removed the call fromLogToolCallto avoid every action being double-counted with a fake 0ms sample alongside the real one.
Changed
pushRecentOCR()signature — Now accepts*OCRResult(with word bounding boxes) instead of juststring, enabling spatial scoping infindRecentOCRBefore.findRecentOCRBefore()signature — Now takestool string, argsJSON stringto enable tool-specific OCR scoping.- Datalog query limit relaxed — From clamped 50–200 range to 1–5000 range (default 50).
[0.2.32] - 2026-07-05
Fixed
chaintool startup panic — shared sub-schema pointers —chainInputSchema()reused the same*jsonschema.Schemaforthen,else, andstepsfields. The MCP SDK’sAddTool()requires schemas to form a tree (not a DAG) and panics on duplicate pointers. Changed to factory functions that return unique instances per call.- Module path mismatch —
go.moddeclaredgithub.com/user/go-mcp-computer-usebut the repo lives atgithub.com/coff33ninja/go-mcp-computer-use. Updated module path and all 7 import references across the codebase. (User was lazy to update this.) - Adaptive engine:
timing_statsandsuccess_ratesnever populated —RecordCommand(which callsRecordResult→RecordTiming+RecordSuccess) was defined but never called. AddedAdaptive.RecordResult(tool, 0, errVal == nil)toLogToolCallso runtime success/failure is tracked per tool. Previouslyagent_analyzealways showed emptytiming_stats: {}andsuccess_rates: {}. - Adaptive engine:
rebuildSequencesmerge corruptsCountand never updatesFreq— When merging duplicate (word, command) entries, Count was overwritten instead of accumulated, and Freq (success ratio) was set at creation and never recalculated on merge. Added internalSuccessCount/FailCountfields toSequenceExample, fixed merge to accumulate correctly and recalculateFreq.
Added
- Chain integration tests — 7 tests build-tagged
//go:build integrationthat start the mcp-server binary and validate chain tool end-to-end via stdio MCP protocol. Covers: simple steps, capture, loop, if/else branching, unknown tool error, timeout, and structured data output. Run withgo test -tags=integration -v -count=1 -timeout 120s ./internal/actions/ -run 'TestChain_'. - CI:
chain-testsjob — runs chain integration tests after lint in.github/workflows/ci.yml. - README badges — Go version, release, CI status, Windows, MCP, last commit, PRs welcome.
[0.2.31] - 2026-07-05
Changed
- All 60+ query/result handlers now return structured data instead of “ok” — Every handler that returns meaningful data (get_volume, get_battery, list_windows, get_system_info, get_uptime, get_clipboard, get_pixel_color, list_displays, get_disk_usage, get_network_info, ocr, find_image, find_all_images, list_audio_devices, list_processes, uia_find, uia_get_text, memory_get/search/list, template_find/list/store/forget, training_, onnx_, datalog_status, chain, launch_and_wait, write_file, delete_file, find_files, list_directory, set/get_working_directory, bridge_debug, set_config, task_begin, and more) now return their structured JSON data instead of the placeholder
"ok"text. The SDK auto-populates tool results from structured output when no explicitTextContentis set, so tools likeget_screen_sizenow show{"width":1920,"height":1080}instead of"ok". get_screen_size,get_cursor_position— same fix applied (noticed during audit).chaintool schema — no longer rejected by Gemini —IfConfig.Then/ElseandLoopConfig.Stepschanged from[]anyto[]ChainStep, which produceditems: trueandtype: ["null", "array"]in the auto-generated JSON schema (both rejected by Gemini’s MCP schema validator). The chain tool now uses a manually craftedInputSchemathat avoids the recursive type cycle injsonschema-goand produces clean schema output.
[0.2.30] - 2026-07-03
Added
- Windows icon embedded in mcp-server.exe — app.ico compiled into a COFF
.sysoresource viarsrc(github.com/akavel/rsrc), so the binary shows a custom icon in File Explorer, taskbar, and title bar. Icon sizes: 16, 32, 48, 64, 256px with SVG source inicons/app.svg. icons/directory — app.svg source, generated PNGs at 5 sizes, app.ico multi-res icon, app.rc resource script.scripts/gen-icons.ps1— PowerShell script that runsrsrcto compileapp.icointocmd/mcp-server/rsrc_windows.syso, auto-installingrsrcif missing. Called frombuild.ps1,lint.ps1, and CI release workflow.
Security
- Bump
golang.org/x/netv0.54.0 → v0.55.0 — dependency update from Dependabot patching a security vulnerability in thego_modulesgroup. .github/dependabot.yml—package-ecosystemwas empty""; set to"gomod"so Dependabot properly scansgo.modfor vulnerabilities.
Changed
write_file—overwriteis now optional — changedOverwrite boolto*Overwrite *boolwithomitempty. No longer required in tool schema, defaults tofalsewhen omitted.get_file_info— returns metadata instead of “ok” — handler now returns actual file info JSON (name,size,is_dir,mod_time,mode).
[0.2.29] - 2026-07-02
Added
- File System tools — 11 new tools for navigating and manipulating files:
list_directory— list files/dirs at path (name, size, is_dir, mod_time, mode)read_file— read file with automatic type detection + native format parsingwrite_file— write/overwrite files with format-aware creation and editingfind_files— recursive glob search (e.g.*.go,**/*.md)copy_file— copy file or directory (recursive)move_file— move/rename file or directorydelete_file— delete file/dir to Recycle Bin via SHFileOperationW (not permanent)create_directory— mkdir -p (recursive directory creation)get_file_info— file/dir metadata (size, mod_time, is_dir, mode)set_working_directory— set working directory for relative path resolutionget_working_directory— get current working directory
- Format-aware
read_file— auto-detects mime type by magic bytes + extension; parses:- Plain text (txt, json, csv, yaml, toml, md, source code, configs, etc.) via
io.ReadAll .docxvianguyenthenguyen/docxlibrary, extracting text from<w:t>XML elements.xlsxviaxuri/excelize/v2, all sheets rendered as TSV.pdfvialedongthuc/pdf, all pages with separators- Images (png, jpg, gif, bmp, tiff, webp) via native WinRT COM OCR (
ocrNative) - Pagination:
pageandpage_sizeparams (default 8000 chars), returns page/totalPages/truncated
- Plain text (txt, json, csv, yaml, toml, md, source code, configs, etc.) via
- Format-aware
write_file— detects target extension and creates/edits:- Plain text — raw write via
os.WriteFile .docx— new file creates from scratch (ZIP+XML); overwrite preserves existing headers/footers/images by swapping<w:body>content, writes to temp + rename.xlsx— new file viaexcelize.NewFile(); overwrite opens existing, repopulates cells from TSV content.pdf— new file creates from text viago-pdf/fpdf; overwrite triespdfcpu.FillFormFilewith JSON form field data, falls back tocreatePdf
- Plain text — raw write via
- File verification —
FilePreCheck/FilePostVerifywired into all 5 action file tool handlers (write, copy, move, delete, create_directory) usingExpConfig/VerifyArgspattern from v0.2.28 - Recycle Bin delete —
delete_fileusesSHFileOperationWwithFOF_ALLOWUNDOto move items to the Recycle Bin instead of permanentos.RemoveAll - Working directory — all file tools resolve relative paths against a configurable working directory (defaults to process CWD, changeable via
set_working_directory) - New dependencies —
nguyenthenguyen/docx,ledongthuc/pdf,xuri/excelize/v2,go-pdf/fpdf,pdfcpu/pdfcpufor native document parsing and creation
Changed
internal/actions/filesystem.go— EnhancedReadFilewith format dispatch + pagination; enhancedWriteFilewith format-aware creation/editing (docx, xlsx, pdf); file verification helpers; Recycle Bin via SHFileOperationW; working directory support.internal/server/server.go— UpdatedReadFileArgs{Path, Page, PageSize}, updatedwrite_filetool description; file verification handlers.internal/actions/chain.go— UpdatedchainReadFilewith page/page_size params.VERSION— bumped to 0.2.29.
Tool Count
Now at 131 total MCP tools (+11).
[0.2.28] - 2026-07-02
Added
- Auto-verify on 5 remaining high-value tools —
open_url,launch_app,find_text_and_click,select_all_and_type,click_menu_itemnow supportauto_verifyandexpectedparameters with OCR-based post-action verification, matching the existing 6 tools. TrainingCatLaunchcategory — New"launch"training category for app launch snapshots.- Pre-action validation (
pre_expected) — All 11 verification-enabled tools now acceptpre_expectedwith the sameExpConfigshape (text/not_text/change). Runs OCR before the action and fails fast if precondition not met — action is never executed. VerifyArgsembeddable struct — Replaced duplicateAutoVerify/Expected/PreExpectedfields across all 11 arg structs with a singleVerifyArgsembed, reducing 22 lines of redundancy.preVerifyCheckhelper — Common pre-verify logic extracted to server.go.- Region-of-typing OCR —
type,type_and_submit,select_all_and_typenow capture cursor position (GetCursorPosition) before typing and restrict verification OCR toSmartRegionAround(cursor, 400px)instead of full-screen scan. - Window-aware OCR for
click_menu_item— Verification scans only within the target window bounds (found byFindWindowByTitle+GetWindowState) instead of full screen. - Coordinate-reuse for
find_text_and_click—FindTextAndClicknow returns click coordinates(int32, int32, error). The handler usesSmartRegionAround(click_pos, 400px)for post-verify instead of full-screen OCR. Pre-verify uses the specified search region.
Changed
internal/actions/chained.go—FindTextAndClicksignature changed fromerror→(int32, int32, error). Callers updated:server.go,chain.go,cmd/benchmark/main.go.internal/actions/training.go— AddedTrainingCatLaunch.internal/server/server.go— All 11 verification handlers updated with cursor-aware region (type tools), window-bounds region (click_menu_item), coordinate-reuse region (find_text_and_click), and pre-verify checks.VerifyArgsstruct replaces repetitive fields.
[0.2.27] - 2026-06-30
Added
find_image/find_all_imagesONNX + OCR fallback — When NCC template matching fails (no match, degenerate template), both tools now cascade through ONNX YOLO object detection → Windows OCR.find_imagereturns the highest-confidence ONNX element (or first OCR word if ONNX is empty).find_all_imagesreturns all ONNX elements + all OCR words. The fallback captures a fresh screenshot if no screenB64 was provided (ensureScreenB64), making it robust against degenerate templates.findImageONNXFallback/findAllONNXFallbackhelpers — Extracted fallback logic with full cascade (ONNX → OCR) and screen capture self-healing.ensureScreenB64helper — Captures screen on demand when the passed-inscreenB64is empty, fixing the edge case where degenerate templates bypass screen decode.ocr_languagestool — NewOcrLanguages()function inocr.go:145queries WinRT COM (IOcrEngineStatics.get_AvailableRecognizerLanguages) and returns every installed OCR language with tag, display name, and native name.- Middle mouse button —
Clicknow supportsbutton: "middle"viamouseEventMiddleDown/mouseEventMiddleUp(0x0020/0x0040). - Horizontal scroll —
Scrollnow takeshorizontal boolparameter; usesmouseEventHWheel(0x1000) flag when horizontal.chainScrollpasses it through from args.
Changed
internal/actions/template.go—FindImageandFindAllImagesrestructured: degenerate templates (zero-dim, no-variance) skip NCC entirely and go straight to ONNX+OCR fallback. Template-larger-than-screen still errors (unrecoverable).FindAllImagesadded as new NCC implementation with non-maximum suppression (overlap >50% suppressed).internal/actions/window_ext.go—GetWindowStatenow returnsfullscreenboolean field. AddedisFullscreen()helper that checksMonitorFromWindow+GetMonitorInfoW+WS_CAPTIONstyle.internal/actions/window_ext.go— AddedmonitorFromWindowproc,MONITORINFOstruct,WS_CAPTIONconstant.docs/tools.md→docs/reference/tools.md— auto-regenerated (2 new tools, 120 total)- Doc audit & reorganization — Audited 20 files across
docs/,dev/,.github/instructions/. Deleteddocs/todo.md(dead) anddev/models-setup.md(duplicate ofdocs/models-setup.md). Deduplicated privacy table, accessibility block, inline tool listing (80 lines), and inline release cycle (8 steps) into cross-references. Reorganized flatdocs/into subdirs:adr/,reference/,guides/,meta/. Moved 15 files to new locations. Updated all cross-references in README.md and all 18 docs files. - CI/CD path updates —
ci.ymlreferencesdocs/tools.md→docs/reference/tools.md;release.ymlreferencesdocs/CHANGELOG.md→docs/meta/CHANGELOG.md;scripts/gen-tools-doc.gowrites todocs/reference/tools.md;scripts/push-and-release.ps1referencesdocs/meta/CHANGELOG.md. scripts/gen-tools-doc.go—MkdirAll("docs")→MkdirAll("docs/reference"), output path changed fromdocs/tools.mdtodocs/reference/tools.md.- New reference docs —
docs/reference/windows-dll-ref.md(every syscall proc, DLL, COM interface used),docs/reference/uipi.md(UIPI elevation detection logic + call sites),docs/reference/com-patterns.md(COM/WinRT patterns: vtable dispatch, async polling, HSTRING/BSTR lifecycle, threading model, UIA tree traversal) - Cross-reference updates —
README.md,codebase-map.md,windows-dll-ref.mdall updated with links to the 3 new reference docs docs/reference/windows-dll-ref.md— Fixed inaccurate advapi32.dll proc names (removedCheckTokenMembership/AllocateAndInitializeSid/FreeSidthat were never used inuipi.go)docs/reference/vtable-verification.md— New doc covering COM vtable stability guarantees, SDK header verification procedure, CI/CD plan, and complete test table (13 tests, 16 unique vtbl indices, all verified)
Fixed
find_image/find_all_imagesno longer error on degenerate templates — Previously a constant-color or 0×0 template returned an unrecoverable error. Now falls through to ONNX + OCR object/text detection.
VTable Verification System (vtable hardening)
- 36 vtblMethod() call sites annotated — Every COM/WinRT vtable dispatch in
uia_com.go,ocr_com.go,ocr.go,winrt.gotagged with inline// N = MethodName (verified 2026-06-30, Win11 26200)and block-level// Verified YYYY-MM-DD — WinBuild SDKheaders above each interface definition. internal/actions/vtable_test.go— 13 smoke tests with build tag//go:build vtable && windows. Covers all 16 unique vtbl indices used in production: UIA GetRootElement (vtbl 5), BoundingRect (43), FindFirst (5, 21), Conditions (21, 23, 25), GetPropValue (10), GetCurrentPattern (16), ValuePattern/Invoke/Array (3, 4, 6, 21), OCR GetLanguages (7, 6), OCR TryCreate (10), OCR TryCreateFromLanguage (9), StorageFile async pipeline (6, 8, 14, 0, 7), HSTRING round-trip, vtblMethod nil guard. All pass (~8s).scripts/verify-vtable-docs.go— Parses all 36 vtblMethod() call sites from source, cross-references each unique index againstcom-patterns.mdandvtable-verification.mddoc tables, checks test coverage via// vtbl:annotations. Rungo run ./scripts/verify-vtable-docs.go. Confirms all 16 unique indices are documented and tested.- CI:
vtable-checkjob —.github/workflows/ci.ymlrunsgo test -v -tags=vtable ./internal/actions/ -run 'TestVtable'after build, plusgo run ./scripts/verify-vtable-docs.goto catch doc drift. go vetpasses — zero warnings across all vtable code.scripts/verify-iid-usage.go— New Go script that parseswinrt.go, scansinternal/actions/for IID references, categorizes each asused(referenced outside winrt.go),internal(only within winrt.go), orunused. Cross-references against the Status column incom-patterns.mddocs. With-update, rewrites the Status column. Run:go run ./scripts/verify-iid-usage.goorgo run ./scripts/verify-iid-usage.go -update.scripts/discover-winrt-iids.ps1-UpdateDocs mode — Added-UpdateDocsswitch that auto-generates the 51-entry IID table indocs/reference/com-patterns.mdfrom discovered values, replacing the content between<!-- IID_TABLE_START -->/<!-- IID_TABLE_END -->markers. Includes fallback pass for statics discovered under flat keys. Now also callsverify-iid-usage.go -updateto set correct usage statuses. CI runs-UpdateDocsthengo run verify-iid-usage.gothengit diffto catch drift.ci.ymlvtable-check job — AddedVerify WinRT IID docs are in sync with sourcestep: runspowershell -File scripts\discover-winrt-iids.ps1 -UpdateDocs+go run ./scripts/verify-iid-usage.go, fails ifdocs/reference/com-patterns.mdhas uncommitted changes.docs/reference/scripts.md— New reference doc cataloging all 8 scripts with purpose, invocation, uniqueness, and cross-references to docs and CI.docs/reference/com-patterns.md— VTable dispatch section (§2) updated withIUnknown basecolumn in all index tables. IID table wrapped in<!-- IID_TABLE_START/END -->markers with new Status column (used/internal/unused). Inline-UpdateDocsflag in both “run the script” instructions.docs/reference/vtable-verification.md— Expanded from “proposed plan” to actual test table withvtbl Indices+Verifiedcolumns, Performance note documenting whyFindAllon desktop root is slow by design, and CI/CD integration reference.uia.goandConditiondead code fixed —CreateAndCondition(vtbl 25) was called frombuildConditionbut not tested; now exercised byTestVtable_IUIAutomation_Conditions.internal/actions/uia_ctrltype_test.go— 3 tests validating the UIA ControlTypeId map covers all 41 constants (50000–50040), no duplicates, no out-of-range values, and unknown names return nil.- Cross-references updated —
README.md,docs/reference/codebase-map.md,docs/reference/windows-dll-ref.md,docs/ci-cd-pipeline.mdall linked to new vtable docs. - VTable indices are hard-won — Every index was researched against SDK headers (UIAutomationClient.h, windows.data.h, windows.storage.h) and verified at runtime on Windows 11 (build 26200). These are the most fragile part of the codebase: an incorrect vtbl index silently calls the wrong method or crashes. The verification script and test suite exist because even a single off-by-one error can take days to diagnose. Microsoft freezes these indices for published interfaces, but when upgrading to a new Windows build, run
go test -tags=vtableandgo run ./scripts/verify-vtable-docs.gobefore shipping.
[0.2.26] - 2026-06-30
Fixed
- Chain tool:
computer_use_*prefix normalization —execToolnow strips thecomputer_use_prefix from tool names before dispatch lookup, so allcomputer_use_*tools (click, type, key_press, ocr, get_screen_size, etc.) work inside chain steps. - Chain
successaggregation —result.Successnow initializes totruebefore checking step results, fixing false-negativesuccess: falsewhen all steps pass. - Keyboard modifier key case sensitivity —
KeyPressnormalizes modifier key names to uppercase beforevkModMap/vkSpecialMaplookups, so"Ctrl","ctrl","CTRL"all correctly match instead of being silently skipped. - Window focus reliability —
FocusWindowusesAttachThreadInputto attach to the target window’s input thread beforeSetForegroundWindow, working around Windows focus-stealing restrictions for background automation processes.
Changed files
internal/actions/chain.go—execTooladdsstrings.TrimPrefix(step.Tool, "computer_use_")at line 312;ExecuteChaininitializesresult.Success = truebefore failure loop at line 200internal/actions/keyboard.go—KeyPressnormalizes keys viastrings.ToUpperbefore modifier/special key lookupsinternal/actions/window.go—FocusWindowusesAttachThreadInputwithGetWindowThreadProcessId/GetCurrentThreadIdinternal/actions/system.go— addedgetCurrentThreadId(kernel32) andattachThreadInput(user32) proc declarations
[0.2.25] - 2026-06-30
Fixed
- Coordinate extraction: case-insensitive key matching —
getIntArgnow falls back to case-insensitive key lookup when an exact match fails, fixing coordinate extraction forclickandmove_mousetools which store their args with capitalizedX/Y(from Gojson.Marshalof struct fields) while the code searched for lowercasex/y. This caused all click coordinate data to be silently ignored byTrainFromDatalog, meaning the__learned__aggregate and per-token coordinate index never accumulated click coordinates.
Changed files
internal/actions/adaptive.go—getIntArgnow does case-insensitive key lookup viastrings.EqualFoldas fallback
[0.2.24] - 2026-06-30
Changed
- Adaptive engine:
__learned__aggregate built from persisted training data —TrainFromDatalognow aggregates all coordinate samples per tool into the__learned__key incoordIndex, so coordinate predictions survive server restart. Combined with the v0.2.23 fallback inpredictCoord,agent_suggestnow returns coordinate predictions forclick/hover/move_mouseusing the aggregate average from all training data, even before any runtime samples accumulate.
Changed files
internal/actions/adaptive.go—TrainFromDatalogstores aggregated coords under__learned__key per tool
[0.2.23] - 2026-06-30
Changed
- Adaptive engine: coord prediction fallback to
__learned__aggregate —predictCoordnow falls back to the runtime-learned__learned__aggregate coordinate when per-token samples incoordIndexare below the threshold of 3. This ensuresagent_suggestreturns coordinate predictions forclick/hover/move_mouseeven when the specific OCR tokens haven’t accumulated 3+ samples yet — the aggregate__learned__accumulates across all invocations of the same tool.
Changed files
internal/actions/adaptive.go—predictCoordnow checks__learned__intoolMapas fallback whentCount < 3, returning the aggregate coord instead ofnil
[0.2.22] - 2026-06-30
Changed
- Adaptive engine: real timing_stats and success_rates —
RecordResultis now called from every action tool’s defer with a captured start time, sotiming_stats(mean, stddev, count, min, max) andsuccess_ratesper tool populate correctly. PreviouslyRecordCommandwas defined but never called, leaving both maps permanently empty.
Changed files
internal/actions/datalog.go— removedAdaptive.RecordResult(tool, 0, ...)fromLogToolCall(moved to per-action defer with real timing)internal/actions/chained.go— addedstartcapture +RecordResulttoLaunchAndWaitandHoverinternal/actions/keyboard.go— addedstartcapture +RecordResulttoKeyDown,KeyUp,KeyPress,TypeTextinternal/actions/mouse.go— addedstartcapture +RecordResulttoClick,MoveMouse,Scroll,Draginternal/actions/window.go— addedtimeimport +startcapture +RecordResulttoFocusWindow
[0.2.21] - 2026-06-30
Changed
- LogToolCall coverage: all 11 MCP action tools now instrumented — Added
LogToolCalltokey_down,key_up,focus_window, andlaunch_and_wait, completing adaptive engine training pair coverage for every non-query action tool. Previously 4 tools produced commands without OCR context pairs, leaving gaps in the training index.
Changed files
internal/actions/keyboard.go— addedLogToolCall("key_down", ...)toKeyDown,LogToolCall("key_up", ...)toKeyUpinternal/actions/window.go— addedLogToolCall("focus_window", ...)toFocusWindowinternal/actions/chained.go— addedLogToolCall("launch_and_wait", ...)toLaunchAndWait
[0.2.20] - 2026-06-30
Changed
- Adaptive engine: OCR bridge auto-complete in
LogToolCall—LogToolCallnow synchronously captures OCR after setting a pending training pair, ensuring every action produces a complete(ocr_before, tool, ocr_after)pair. Previously pairs only completed when the next explicitOCRScreen()call happened, causing all training sequences to cluster under “click”. Also addedLogToolCalltoHoverandMoveMousewhich were missing it entirely.
Changed files
internal/actions/datalog.go—LogToolCallauto-captures OCR after pending pair setinternal/actions/chained.go— addedLogToolCall("hover", ...)toHoverinternal/actions/mouse.go— addedLogToolCall("move_mouse", ...)toMoveMouse
[0.2.19] - 2026-06-30
Changed
- Keylogger rewrite: hooks → polling — Replaced
WH_MOUSE_LL+WH_KEYBOARD_LLlow-level hooks withGetAsyncKeyStatepolling loop (50ms ticker). Eliminates the system-wide input lag caused by the Go hook callback trampoline on every mouse event. The polling loop runs in a goroutine with no locked OS thread and no Windows message loop. Trade-off: scroll wheel events no longer detectable (acceptable cost for eliminating system-wide input lag).
Fixed
-
CI lint failure — stale tools.md & uncategorized tools —
scripts/gen-tools-doc.gowas missing category entries for 4 tools (bridge_debug,introspection_analyze,task_begin,task_end), causing them to fall under “Uncategorized” anddocs/tools.mdto show 114 instead of 118 tools. The lint check (regenerate + diff) then failed, skipping the build job. Added"Introspection & Debugging"category, removed staledocs2/staging output from the script, and regenerateddocs/tools.md. -
yolo_datasetlocation inconsistency — removed staleyolo_dataset/from repo root (empty train/val dirs).export_yolo_datasetnow defaults to%APPDATA%\go-mcp-computer-use\yolo_dataset\whenoutput_diris omitted. Addedyolo_dataset/to.gitignoreto prevent future repo root drift.
[0.2.18] - 2026-06-29
Added
- Post-Task Introspection Engine (
internal/actions/introspection.go) — three new MCP tools for task-aware self-improvement:task_begin— marks task start with description, timestampstask_end— closes task, mines insights from command_log between start/end: slowest tools, most failed tools, OCR stats, repeated command patterns, and improvement suggestionsintrospection_analyze— browse completed task history with full insight data- Uses existing
command_log+ocr_logtables — no new logging infra needed task_logtable added to datalog DB
Changed
datalog_statusnow reportstask_countin stats
[0.2.17] - 2026-06-29
Fixed
- OCR→Training bridge window —
bridgeWindowincreased from 3s to 30s. The OCR→AI→MCP→Click round trip regularly exceeded the original 3-second window, preventing training pair creation. Debugged via newbridgeBufferSize()andBridgeDebugInfo()diagnostic functions exposed through thebridge_debugMCP tool.
Added
bridge_debugMCP tool — debug the OCR→command bridge state, showing recent OCR buffer contents, pending command, and timing info.
[0.2.16] - 2026-06-29
Added
- Adaptive Engine (
internal/actions/adaptive.go) — pure Go statistical ML system with three components:- TimingTracker — rolling-window (N=100) per-tool statistics: mean, stddev, min, max. Auto-suggests adaptive delays based on historical execution time plus success-rate multiplier (1.5× by default, 3× when success rate < 50%).
- SuccessTracker — per-tool success/failure ratios. Queried on every
SuggestDelay()call to adjust timeouts. - SequencePredictor — TF-IDF-style word index from
training_pairs. Given OCR text, tokenizes and scores each word→command mapping by historical success frequency. Returns ranked predictions with confidence (0.0–1.0) and sample size.
- MCP Resources (5) — auto-exposed to the AI client, read on every session context:
datalog://stats— current row counts for all four datalog tablesdatalog://commands— 20 most recent command log entriesdatalog://ocr— 10 most recent OCR snapshotsdatalog://pairs— 20 most recent training pairsadaptive://analysis— full adaptive engine analysis (timing stats, success rates, learned sequences)
- Agent MCP Tools (3) — AI-queryable loop for context-aware decisions:
agent_analyze— returns full timing stats, success rates, and top learned sequences for AI decision-makingagent_suggest— given OCR screen text, predicts the best next command ranked by confidenceagent_train— rebuilds the word→command index from currenttraining_pairstable
- Auto training pair generation — passive OCR bridge creates triple (ocr_before, command, ocr_after) without slowing commands:
- Ring buffer of last 5 OCR snapshots with timestamps
- Every command auto-pairs with most recent OCR (within 3s window) as
ocr_before - Next OCR snapshot completes as
ocr_after - Command stored as
{"tool":"name","args":"..."}JSON for robust parsing
Fixed
datalog_querytable name mismatch — switch-case expected short names ("commands","ocr","chains","pairs") but the handler passed raw table names. Now accepts both forms as aliases.TrainFromDatalogJSON parsing — robustextractToolFromJSONhelper handles both JSON{"tool":"..."}and plain string command values.
Changed
- Tool count — 111 → 114
- VERSION — bumped 0.2.15 → 0.2.16
- gen-tools-doc.go — added “Adaptive Agent” category
- LogCommand — now releases SQLite lock before OCR bridge to avoid deadlock with LogOCRSnapshot (no cross-lock ordering)
- LogTrainingPair —
Commandfield stores structured{"tool":"name","args":"..."}JSON instead of raw args string
Documentation
- docs/tools.md — regenerated with 114 tools across categories including “Adaptive Agent”
[0.2.15] - 2026-06-29
Added
- Data logging database (
internal/actions/datalog.go) — new SQLite DB at%APPDATA%/go-mcp-computer-use/datalog/datalog.dbwith four tables:command_log— every chain/tool execution with args, success, duration, error textchain_log— full chain executions with step counts, success/fail breakdown, chain JSONocr_log— OCR snapshots with full OCR text, word count, linked screenshot image pathtraining_pairs— OCR-before + command + OCR-after triples for ML sequence learning
-
Automatic logging hooks — chains, individual commands, and OCR calls are logged automatically via goroutines with no performance impact on the main execution path.
- Three new MCP tools:
datalog_query— query any table (commands, chains, ocr, pairs) with filters (source, tool, success), returns rows as JSONdatalog_export— export training pairs as JSON array for downstream ML training pipelinesdatalog_status— get row counts for all four tables
Changed
- VERSION — bumped 0.2.14 → 0.2.15
- Tool count — 108 → 111
[0.2.14] - 2026-06-29
Added
-
NormalizedElementcoordinate system — element positions stored as window-relative 0.0–1.0 fractions viaWindowNormalizerininternal/actions/dpi.go. Layout-independent across screen resolutions and multi-monitor. IncludesGetDPIScaleForWindow,Normalize/Denormalizehelpers, andProportionalRegionfor computing screen-absolute OCR crops as a percentage of the active window. -
OCRProportionalWindowRegion— new OCR function inocr.gothat takes a window handle + proportional fractions, eliminating hardcoded pixel crops. -
Auto-expand tiny OCR regions —
FindTextAndClicknow detects crops <300px in any dimension and falls back to a generous 5%–95% of the active window. Prevents “Desktop not found” failures on small fixed-pixel regions. -
Window context in ONNX detection —
DetectionOutputcarriesWindowTitleandNormalized []NormalizedElementalongside absolute coordinates. Computed per-active-window during inference. -
Training schema migration —
training_samplestable gainswindow_rect TEXTandnormalized_coords TEXTcolumns.saveTrainingSampleDirectaccepts and persists both normalized coords and window rect JSON.
Fixed
NormalizeElementmissing Class/Confidence copy —WindowNormalizer.NormalizeElementreturned aNormalizedElementwith zeroedClassandConfidencefields. Exposed by round-trip test (TestNormalizeElementRoundTrip). Now copies both fields before returning.
Changed
- Watcher cache —
CachedDetectionincludesNormalizedelements alongside absoluteElements. Training samples from watcher snaps now carry window rect context. - VERSION — bumped 0.2.13 → 0.2.14
Tests
- Coordinate system tests —
dpi_test.gowith 6 tests covering: normalize/denormalize round-trip, coordinate bounds (corners, center, size), proportional region math,NormalizeElementclass/confidence round-trip, and zero-size window edge case.
[0.2.13] - 2026-06-29
Fixed
- ONNX detection timeout (65s → 599ms) — root cause was not DLL incompatibility but performance:
parseYOLOOutputpassed all 8400 raw detections through NMS at O(n²) = ~15M iterationsMemoryStoreDetectionElementscalledMemorySet5507 times — each a separate SQLite INSERT with global mutex lock- Fixed:
parseYOLOOutputnow applies confidence threshold early (0.25), pre-filtering to ~50 boxes before NMS - Fixed:
MemoryStoreDetectionElementsrewritten with batched SQLite inserts in a single transaction, capped at 200 elements
Changed
- ONNX Runtime DLL updated — v1.20.1 → v1.26.0 to support opset 22 (required by yolo11n.onnx). Limited opset support warning is non-fatal.
[0.2.12] - 2026-06-29
Fixed
- Release binaries crash with STATUS_ILLEGAL_INSTRUCTION — Zig cc on GHA runners defaults to
-march=native, generating CPU-specific instructions incompatible with older machines (Pentium Gold G5400). Pinned-mcpu=x86_64_v2inCGO_CFLAGSso binaries run on any x86-64 CPU. - CGO_LDFLAGS also needs
-mcpu=x86_64_v2—actions/setup-go@v5overridesCGO_LDFLAGSwith-O2 -g, dropping the CPU baseline. BothCGO_CFLAGS(compile) andCGO_LDFLAGS(link) now pin-mcpu=x86_64_v2.
Changed
scripts/build.ps1— addedCGO_CFLAGSwith-mcpu=x86_64_v2baseline for portable builds.github/workflows/release.yml— same CPU baseline pin in bothCGO_CFLAGSandCGO_LDFLAGS, plus-fno-sanitize=alland-Wno-error
[0.2.11] - 2026-06-29
Added
scripts/gen-tools-doc.go— parsesinternal/server/server.goformcp.AddToolcalls, generatesdocs/tools.mdwith categorized 108-tool listing. CI validates freshness on every push/PR.scripts/push-and-release.ps1— one-shot auto-release: reads VERSION, commits with changelog body, tags, pushes, waits for release workflow, downloads binary, replacesmcp-server.exe, restarts OpenCode Desktop as admin.docs/tools.md— auto-generated tool reference doc (never stale).docs/security.md,docs/configuration.md,docs/build.md,docs/architecture.md,docs/accessibility.md— split from monolithic README.- Weekly module maintenance —
.github/workflows/mod-maintenance.ymlrunsgo get -u ./...+ auto-PR every Monday. - CI:
go mod tidyvalidation — fails ifgo.mod/go.sumdrifts from tidy state.
Changed
- README.md — collapsed 383→92 lines, links to focused docs/ split.
- Root docs moved —
plan.md,todo.md,backlog.md,known-issues.md,CHANGELOG.mdrelocated todocs/. - CGO mandatory — removed all
-NoCGOflags, pure-Go fallback paths, and optional-CGO language across 9 files.release.ymlnow produces a singlemcp-server.exe(CGO+Zig). - Release workflow — single binary output, no
-cgosuffix variant. scripts/build.ps1— removed-NoCGOswitch, always requires Zig cc.
Documentation
- README split — large sections moved into focused docs for maintainability.
- All NoCGO references removed — across
plan.md,adr-002,comparison-vs-alternatives.md (formerly comparison-vs-windows-recall.md),ci-cd-pipeline.md,build.md,README.md.
[0.2.10] - 2026-06-29
Documentation
- Systematic doc audit — fixed 90 stale statements across 12 docs: tool counts (103→108 restored from actual registrations), version refs, CGO/dependency claims, category counts, missing tool listings, completed Slice 4 checkboxes, stale future-tool lists
- Architecture guide — added Part 6 to computer-use-guide: layered agent stack (LLM→MCP→Controller→Perception→Memory→World), ML vision + spatial memory, division of responsibilities, convergence of LLM+MCP+ML
- Source fix — server.go tool count hardcode corrected 103→108 to match actual registrations
- Config auto-start — watcher_auto_start config created on dev machine
Changed
- VERSION — bumped 0.2.9 → 0.2.10
[0.2.9] - 2026-06-29
Added
scripts/build.ps1— unified build script with-UseZigflag for CGO-enabled builds- CI/CD: CGO + Zig cc build pipeline — CI now runs two jobs: no-CGO lint+build and CGO+Zig build. Release workflow produces both
mcp-server.exe(no CGO) andmcp-server-cgo.exe(with ONNX support). - Zig 0.16.0 support —
scripts/install.ps1updated to download Zig 0.16.0
Documentation
- README.md — documented CGO requirements for ONNX tools with Zig cc build instructions
- known-issues.md — B13: ONNX tools require CGO (documented workaround)
- Tool count docs updated — all docs updated to 108 tools, stale CGO claims corrected
[0.2.8] - 2026-06-29
Added
-
key_down/key_upMCP tools — separate key hold/release for game-play sequences. Chains can now hold movement keys while dragging camera and pressing abilities, all server-side with no round-trip latency.KeyDown("W")holds the key,KeyUp("W")releases. Full VK support including modifiers, letters, digits, and special keys. -
keylogger_start/keylogger_stop/keylogger_statusMCP tools — record real keyboard + mouse input (keys, clicks, drags, moves, scroll) via low-level Windows hooks (WH_KEYBOARD_LL+WH_MOUSE_LL). Output is a chain-compatible JSON sequence for AI replay. Includes timing-accurate delays between events. Mouse clicks auto-detect drag vs click by distance/time thresholds. Mouse moves throttled to meaningful position changes. -
sendVKPresshelper with 50ms inter-key delay —KeyPress,TypeText,sendCharWithVKnow usesendVKPress(vk)which inserts a 50mstime.Sleepbetween key down and key up. Fixes game engines and DirectInput applications that miss instant down/up sequences (character switch hotkeys 1-4, ability keys).
Fixed
-
warnElevatedfalse positive when both server and target are elevated —warnElevated()only checked if the foreground window was elevated, not the MCP server itself. If both are elevated (server running as Admin targeting an admin game),SendInputkeyboard works fine, but the check falsely blocked it. AddedisSelfElevated()— only blocks keyboard when server is non-elevated AND target is elevated. -
KeyPressmodifier ordering —["CTRL", "C"]sentCvia Unicode first, then pressed Ctrl down, then released Ctrl. The key arrived before the modifier was held. Rewrote to process keys in order: modifiers are pressed immediately, target keys are sent while held, all modifiers released in reverse at end. - Keyboard input uses VK codes instead of
KEYEVENTF_UNICODE—KEYEVENTF_UNICODEsynthesizesWM_CHARmessages, which many applications ignore (game engines, terminals, code editors, browser input fields). Rewrote all keyboard functions to use VK codes:TypeTextandTypeAndSubmitusesendCharWithVK()— maps each character to its VK code + Shift state usingcharToVKtable (US keyboard layout). Letters, digits, punctuation all handled.KeyPresssends all keys (letters, digits, special keys) as VK codes. Modifier combos like Ctrl+C now work correctly: Ctrl down → VK_C → Ctrl up.
Dragincremental movement — was sending a single jump from start to end (mouseDown → teleport → mouseUp). Games and map UIs ignored this as a teleport. Now sends 5–50 incremental steps with 5ms delays, proportional to distance. Map panning now works correctly.
Changed
sendUnicoderemoved — no longer used. All keyboard input via VK codes.- Tool count: 103 → 108 (added
key_down,key_up,keylogger_start,keylogger_stop,keylogger_status).
Documentation
- Elevation & UIPI section in README — explains admin vs non-admin behavior
- Known issues B11, B12 — documented keyboard issues and fixes
[0.2.7] - 2026-06-29
Added
- Statistical prior model (
priors_statstool) — Go-native “training” without Python. Element frequency + position distributions are learned per window from collected training samples. Priors boost confidence for expected elements (e.g., “laptop” in browser windows) and suppress unlikely ones (e.g., “tv” in code editor). Position outliers beyond 3σ are penalized. - Prior-based confidence adjustment —
ONNXDetectnow callsAdjustConfidenceWithPriors()after NMS, adjusting every detection’s confidence based on learned per-window statistics. Gated byprior_adjustmentconfig field (default:true). export_yolo_datasettool — exports unused training samples (signal_level >= 1) as a YOLO-format dataset (images + normalized label files + train/val split +dataset.yaml). Users with Python can train externally via Ultralytics.training_cleanup_noisetool — deletes low-signal (signal_level=0) samples older than a threshold. Supportsdry_run=trueto preview deletions. Frees disk space from watcher noise frames.training_enabledconfig field — when set tofalse, disables all auto-save training snapshots (both from actions and the background watcher). Default:true.prior_adjustmentconfig field — when set tofalse, disables prior-based confidence adjustment in ONNXDetect. Default:true.- Priors auto-populated on save — every training sample save (raw or watcher) also updates element priors via
UpdatePriorsFromDetections. Negative samples (zero elements) update frequency denominators.
Changed
set_configtool — runtime config changes without restart. Accepts:training_enabled(stop/start background data collection),prior_adjustment,verify_bounds,log_level,watcher_enabled(start/stop watcher),watcher_interval_seconds(change polling frequency live). Changes persist to disk immediately. Enables users to disable data collection and control the watcher mid-session for privacy or debugging.watcher_auto_start/watcher_interval_secondsconfig —watcher_auto_start: truestarts the background watcher on server boot with the configured interval. Default:false.- Tool count: 99 → 103 (added
priors_stats,export_yolo_dataset,training_cleanup_noise,set_config).
Fixed
SendInputsilently dropping mouse clicks — theinputstruct inmouse.gohad an orphan_ [8]bytepadding field, makingunsafe.Sizeof= 48 bytes. Windowssizeof(INPUT)on x64 is 40 bytes.SendInputreturns 0 whencbSizedoesn’t match, soSetCursorPosmoved the cursor but the click event never fired. Removed the extra padding — struct is now exactly 40 bytes.- Network struct layout mismatches —
IP_ADDR_STRINGwas missing_ [4]bytetrailing padding (44→48 bytes).IP_ADAPTER_INFOandFIXED_INFOused[260/132]uint16forchararrays (2x Windows size, shifting every subsequent field). Changed to[260/132]byteand added alignment padding afterDhcpEnabled. - All Windows API structs verified — audited every struct passed to Win32 via
unsafe.Pointerininternal/actions/:mouseInput(32B ✓),input(40B ✓),point(8B ✓),keyboardInput(24B ✓),inputKbd(40B ✓),BITMAPINFOHEADER(40B ✓),RECT,MONITORINFOEXW,DEVMODEW,MEMORYSTATUSEX,SYSTEM_POWER_STATUS,PROCESSENTRY32W,LASTINPUTINFO,VARIANT,UiaRect,WinRect— all match Windows x64 sizes.
Changed
Dragrewritten for raw input games — replacedSetCursorPos(invisible to DirectInput/raw input) withSendInput+MOUSEEVENTF_MOVE | MOUSEEVENTF_ABSOLUTE. Coordinates normalized to 0–65535 range. Game engines using raw input now see the movement between mouse-down and mouse-up.
Documentation
- Elevation & UIPI section added to README — explains admin vs non-admin behavior (keyboard warns, mouse silently fails), how to run elevated, and reassurance that normal apps work fine without elevation.
[0.2.6] - 2026-06-28
Added
- Training data pipeline (
training_save_sample,training_list_samples,training_stats,training_mark_used) — persistent screenshot + ONNX detection storage for model fine-tuning. Images saved to categorized folders (raw/click/,raw/type/,raw/navigate/,raw/ocr/,raw/general/,watcher/elements_found/,watcher/no_elements/) with metadata insamples.db. Each sample carries atask_promptstring that the ML learns to predict during training. - Auto-save on every UI action —
click,type,scroll,drag,hover,key_press,type_and_submit,select_all_and_type,browser_navigate,browser_search,open_url, andfind_text_and_clickhandlers (both direct MCP and chain steps) automatically capture a screenshot + ONNX detection + save toraw/{category}/with the action description astask_prompt. find_ui_elementtool — three-layer cascading element locator: checks memory first (cached ONNX detections by window+label), runs ONNX detection with label matching, falls back to OCR for text elements. Stores findings in memory for reuse. Saves training samples (positive + negative).- Memory-backed element caching — every
ONNXDetectcall auto-stores detected elements as memory facts (memory_set, scopeui, keyedui:{window_title}:{class}) with 1-hour TTL. AI can query memory for known element locations without re-running ML. - Quality/signal filtering — every training sample gets a
signal_level(0=noise, 1=elements found, 2=elements+task context).training_list_samplesacceptsmin_signalfilter. Noise samples (watcher frames with zero elements) are flagged for discard.
Changed
- Restructured training directories — from flat
samples/{cat}_{ts}.pngtoraw/{cat}/{ts}.png+watcher/{cat}/{ts}.pnglayout. Database renamed fromtraining.dbtosamples.db. - Watcher save path — frames now save to
watcher/elements_found/orwatcher/no_elements/instead of flatreferences/dir. - ONNXDetect no longer auto-saves — removal of inline
saveTrainingSampleDirectin ONNXDetect to avoid caller confusion. Watcher handles persistence; explicit calls handle the rest.
[0.2.5] - 2026-06-28
Fixed
memory_setschema validation —MemorySetArgs.Value anygenerated"value": truein JSON Schema, which OpenCode’s MCP validator rejected. Fixed with explicitInputSchemausingjson.RawMessage+ description-only schema.close_windowWin32 API — was callingShowWindowAsync(hwnd, 0x10)but0x10 = WM_CLOSEis not aShowWindowcommand. Changed toPostMessageW(hwnd, WM_CLOSE, 0, 0).onnx_statusglobal state bug — used globalmodelsDirwhich was empty whenInitONNXfailed. Now callsgetModelsDir()directly.
Added
- Background watcher (
onnx_watch_start/stop/status/cache) — goroutine that periodically captures screen, runs ONNX detection, caches last 20 results, and auto-saves reference PNGs when detection returns zero elements. savePNGauto-save in detection —onnx_detectnow saves aref_<ts>.pngto%APPDATA%/go-mcp-computer-use/models/references/when detection returns zero elements (AI confusion signal).focus_window_by_title— finds window by title, focuses, and clicks title bar to ensure activation.- Browser automation —
browser_focus_url_bar,browser_new_tab,browser_navigate,browser_search. - File Explorer automation —
explorer_focus,explorer_open_path. uia_warmupconfig field and async UIA warmup on startup.
Changed
- Eliminated Python dependency entirely — removed
convertYoloToONNX(),detectWithPython(),pythonDetectResultstruct,os/exec,bytes,stringsimports. - Switched YOLO model — from HuggingFace
best.pt(PyTorch, 57 MB, 7 UI classes) to Ultralytics pre-exportedyolo11n.onnx(10.9 MB, 80 COCO classes).
[0.2.0] - 2026-06-27
Changed
- v0.2.x branch baseline — cut from v0.1.11 as starting point for v0.2 development. All subsequent changes on this branch increment as
+0.0.1(v0.2.1, v0.2.2, etc.).
[0.2.1] - 2026-06-27
Added
chaintool — sequential step executor that runs multiple tools server-side without round trips. Supportstool(call any registered tool),wait(sleep N ms), andcapture(save step output as `` for use in subsequent steps). Error modes:stop(halt on first error, default) orskip. Global timeout. 40+ tools dispatched.- Variable substitution — `` in string args is replaced with captured output from earlier steps.
- ChainFromJSON — convenience entry point for programmatic chain execution from JSON string.
[0.2.4] - 2026-06-28
Added
memory_set/memory_get/memory_search/memory_list/memory_forgettools — SQLite-backed memory store usingmodernc.org/sqlite(pure Go, zero CGO). Database at%APPDATA%/go-mcp-computer-use/memory.dbwith WAL mode, FTS5 full-text search, auto-syncing triggers, TTL support, scope isolation, and tag filtering.layout_validatetool — validates stored UI element layouts against the current screen. Checks window existence, position drift (with tolerance), and OCR keyword verification around element coordinates. Returns per-element confidence (ok/drifted/stale) with adjusted coordinates.template_store/template_find/template_list/template_forgettools — self-growing template library.template_storeauto-crops a 48×48 PNG template around a coordinate from the current screen and stores it in theelement_templatestable.template_finduses NCC template matching (find_image) to relocate the element visually on the current screen, returning coordinates and drift. Hit count auto-increments on each successful find, enabling the system to self-train over time.onnx_status/onnx_detect/onnx_downloadtools — ONNX ML backend for UI element detection.onnx_statuschecks runtime and model availability.onnx_detectruns YOLO11s inference on a screenshot or full screen to detect UI elements (button, textbox, checkbox, dropdown, icon, tab, menu_item) with bounding boxes and confidence scores. Usesgithub.com/yalue/onnxruntime_gofor native ONNX Runtime support. Requires manual download ofonnxruntime.dlland model files. Falls back gracefully when runtime/models are missing.focus_window_by_titletool — focus management for reliable keyboard input. Finds a window by title, focuses it, and clicks its title bar to ensure activation.ChainStep.FocusWindowfield — chain steps can specifyfocus_window: "window title"to auto-focus and activate the window before executing the step. The chain executor handles window lookup, focus, title bar click, then runs the step.browser_focus_url_bar/browser_new_tab/browser_navigate/browser_searchtools — generic browser automation (Firefox, Chrome, Edge, Brave, Opera).browser_focus_url_barfocuses the URL bar (Ctrl+T for Firefox, Ctrl+L for others).browser_new_tabopens a new tab (Ctrl+T).browser_navigateopens a new tab and navigates to a URL.browser_searchopens a new tab and performs a search query. Backed byBrowserFocusURLBar,BrowserNewTab,BrowserNavigate,BrowserSearchininternal/actions/browseruse.go— reusable composite functions that import existing modules instead of duplicating logic.explorer_focus/explorer_open_pathtools — File Explorer automation.explorer_focusfinds and activates an existing File Explorer window by title.explorer_open_pathopens explorer at a given path, reusing existing windows when possible (Ctrl+L + path) or launching a new one. Backed byExplorerFocus,ExplorerOpenPath,ExplorerNavigateToininternal/actions/windowexploreruse.go.
Changed
- Replaced
firefox_focus_url_bar— removed Firefox-specific function fromchained.go. Replaced with genericbrowseruse.gothat detects browser type from window title and uses browser-specific keyboard shortcuts (Ctrl+T for Firefox URL bar, Ctrl+L for Chrome/Edge). - Refactored
FocusWindowByTitle— now delegates to sharedfocusAndActivateWindowhelper, reducing duplication across browser, explorer, and generic focus code paths.
Removed
FirefoxFocusURLBar— removed frominternal/actions/chained.go. Superseded byBrowserFocusURLBar. Tool name changed fromfirefox_focus_url_bartobrowser_focus_url_bar.
[0.2.3] - 2026-06-28
Fixed
TypeAndSubmitEnter viaKeyPress— appended\rusedsendUnicode(0x0D)which sends the CR character viaKEYEVENTF_UNICODE, unreliable in Firefox/browser address bars. Replaced withKeyPress([]string{"ENTER"})with a 50ms pause, matching the same code path used by thekey_presshandler.
[0.2.2] - 2026-06-28
Added
pollstep type — polls OCR atevery_msinterval untilocr_containstext is found ortimeout_mselapses. Syntax:{"poll": {"every_ms": 1000, "timeout_ms": 30000, "ocr_contains": "Submit"}}.ifstep type — OCR checks forocr_containstext, executesthenorelsebranch. Syntax:{"if": {"ocr_contains": "Error", "then": [...], "else": [...]}}.loopstep type — repeats sub-stepstimesiterations. Syntax:{"loop": {"times": 5, "steps": [...]}}.StepResult.Steps— nested step results for if/loop sub-steps, visible in chain output.- UIA warmup at server startup — pre-initializes COM and creates/releases a UIA instance, absorbing the one-time 15-42s cold-start cost so handlers respond instantly.
WarmupUIA()— exported function to pre-warm COM/UIA at server startup.
Fixed
- StepResult.Index always
0—execWait/execToolcreated freshStepResultstructs discarding the loop index. Index is now set after the switch. SelectAllAndTypeuses VK codes —sendUnicode(0x01)usedKEYEVENTF_UNICODE(VK_PACKET) which doesn’t trigger select-all in most apps. Replaced withsendVK(VK_CONTROL)+sendVK(VK_A)for reliable Ctrl+A.- Variable substitution supports dotted paths — regex
[a-zA-Z0-9_]+didn’t match ``. Updated to[a-zA-Z0-9_.]+withresolveVarPath()for nested map lookups. SelectAllAndTypeelevated warning — now callswarnElevated()before sending input, preventing silent drops on admin windows.
[0.1.11] - 2026-06-27
Added
- VERSION file + ldflags — single source of truth at project root, injected via
-X main.Version, replaces hardcoded string - CI/CD pipeline —
.github/workflows/ci.yml(build + vet on push/PR),.github/workflows/release.yml(tag-triggered GitHub Release with binary + SHA256 + changelog) .govetallow— documents COM/WinRT unsafe.Pointer conventions for vet policyscripts/lint.ps1— local CI runner: vet + build + tests
Changed
- COM types — all interface pointers stored as
unsafe.Pointerinstead ofuintptr:uiaAuto.p,uiaCondition.p,uiaElement.p,uiaElementArray.p,bstrToGoparameter,getCurrentPatternreturn type vtblMethod— rewritten withunsafe.Pointerparameter +unsafe.Add, satisfies vet’s unsafeptr checker- Syscall output params — all local variables receiving COM pointers via SyscallN declared as
unsafe.Pointerinstead ofuintptr - GUID literals — all 14
windows.GUIDvalues inwinrt.gouse keyed fields - CI workflows — use
scripts/lint.ps1instead of rawgo vet
[0.1.10] - 2026-06-27
Fixed
- Keyboard VK-coded keys (Enter, Backspace, Tab, Ctrl+letter) sent via
sendKey/KeyPresswere silently dropped by the system — onlyKEYEVENTF_UNICODEpath worked. Rewrote keyboard handling to send all keys throughKEYEVENTF_UNICODEwhere possible: special keys map to Unicode control characters (Enter=0x0D, Backspace=0x08, Ctrl+A-Z=0x01-0x1A). VK fallback only for non-printable keys (arrows, F-keys, Insert, etc.) TypeAndSubmitandSelectAllAndTypenow use Unicode path instead of VK-codedKeyPressfor Enter and Ctrl+A
[0.1.9] - 2026-06-27
Added
- B9: UIPI elevation detection for keyboard input (
TypeText,KeyPress) — returns clear warning when foreground window is elevated (admin), instead of silently dropping input
[0.1.8] - 2026-06-27
Fixed
- B3:
list_displaysonly returned primary monitor —monitorEnumProcgated onMONITORINFOF_PRIMARYflag, skipping all non-primary displays
[0.1.7] - 2026-06-27
Fixed
- B4:
uia_get_text/uia_invokeno longer crash MCP transport —GetCurrentPatternnil check added before pattern operations
[0.1.6] - 2026-06-27
Fixed
- B2:
list_audio_devicesreturns[]instead ofnull— empty PowerShell output produced nil slice which serialized as JSONnull
[0.1.5] - 2026-06-27
Fixed
- B6:
Wait()calculation was 1 million times too long —NtDelayExecutionargument was-(ns * 10000)instead of-(ms * 10000), causinghover(and any tool callingWait) to block for hours instead of milliseconds
[0.1.4] - 2026-06-27
Fixed
- B1:
get_brightnessreturns clear “brightness not supported on this display” instead of parse error when display doesn’t support WMI brightness control (desktop monitors)
[0.1.3] - 2026-06-27
Fixed
- B5:
screenshot_elementnow clamps off-screen window coordinates to screen bounds instead of rejecting them (e.g., windows withx=-8from Aero Snap) - Multi-monitor:
ScreenSize()now returns virtual desktop dimensions (SM_CXVIRTUALSCREEN/SM_CYVIRTUALSCREEN) instead of primary monitor only, fixing coordinate validation across multiple displays
[0.1.2] - 2026-06-27
Fixed
- UIA COM and OCR WinRT apartment model conflict: changed
RoInitializefromRO_INIT_SINGLETHREADEDtoRO_INIT_MULTITHREADEDso both UIA and OCR use MTA on the same thread, preventingRPC_E_CHANGED_MODEerror
[0.1.1] - 2026-06-27
Added
- Native WinRT COM OCR — replaces PowerShell OCR with direct COM calls:
StorageFile.GetFileFromPathAsync→OpenAsync→BitmapDecoder.CreateAsync→GetSoftwareBitmapAsync→OcrEngine.RecognizeAsync. Zero CGO, no Windows SDK needed. - Native COM UI Automation — replaced PowerShell UIA with direct COM calls to
UIAutomationCore.dll(IUIAutomation, IUIAutomationElement, conditions, patterns). All operations via native COM. - WinRT COM infrastructure (
winrt.go) — HSTRING management,RoInitialize,RoGetActivationFactory,IAsyncInfopolling, COM helpers - OCR falls back to PowerShell if native COM fails
Changed
- All OCR and UIA operations now use native COM instead of PowerShell — 2-8x faster
- OCR full screen: 653→292ms (2.2x)
- OCR region 400×400: 542→68ms (8x)
- find_text_and_click: 809→275ms (2.9x)
comReleasesignature changed fromuintptrtounsafe.Pointerfor unified COM cleanup- ADR-002 updated: project now uses native COM/WinRT, not just Win32 API
Fixed
- WindowsGetStringRawBuffer signature: actual DLL export returns buffer pointer in RAX (2 params), not as out parameter (3 params) — MSDN docs differ from Win10 10.0.26100 behavior
- All vtable reads: corrected
*(*[N]uintptr)(obj)pattern (reads object data) tovtblMethod()(reads actual vtable entries) - OCR PowerShell script: properly loads WinRT types via
WindowsRuntimeSystemExtensions.GetAwaiterwithMakeGenericMethod, fixing OCR on systems where WinRT async extension methods don’t resolve in PowerShell 5.1 - Go raw string literal: avoids backtick in
IAsyncOperation1by using-like` wildcard matching
[0.1.0] - 2026-06-27
Added
- Screenshot (full + region) via GDI BitBlt
- Mouse control: click, move, scroll, drag, hover
- Keyboard input: type, key_press, type_and_submit, select_all_and_type
- OCR via Windows.Media.Ocr with language support
- Template matching via normalized cross-correlation
- find_text_and_click, wait_for_text, click_menu_item, launch_and_wait
- Screen recording (duration_ms + interval_ms → base64 frames)
- Window management: list, focus, find, move, resize, minimize, maximize, restore, close, get_state
- Audio devices: list playback/recording, set default
- Clipboard: get/set with retry + timeout
- System: volume, mute, brightness, battery, disk, DPI, display info, uptime, idle
- Network: hostname, IPs, DNS, gateway, ping
- Processes: list, launch, kill
- Power: shutdown, restart, sleep, hibernate
- Per-monitor DPI awareness
- UI Automation via PowerShell: find elements, get text, invoke
- get_display_modes tool (69th tool) — enumerate all display modes
- Config file:
~/.config/go-mcp-computer-use/config.json - Install script:
scripts/install.ps1with Zig cc support
Changed
- syscall hardening:
ptr()helper for safe unsafe.Pointer conversion - performance optimizations across all action modules
- README with comprehensive tool listing and security warning
- MCP client configs documentation for 19 agents
Security
- Added SECURITY WARNING section to README detailing dangerous capabilities