---
name: computer-use-linux
description: "Use when computer_use fails on Linux X11 or SSH sessions."
version: 1.0.0
author: Hermes Agent
license: MIT
platforms: [linux]
metadata:
  hermes:
    tags: [computer-use, cua-driver, x11, display, ssh, vision, linux]
---

# Computer Use on Linux (X11 / SSH)

Getting `computer_use` (cua-driver) to actually see and drive a Linux desktop,
especially when Hermes runs inside an SSH/TUI session that has no graphical
environment of its own.

## When to Use

- `computer_use` capture returns `0x0` / empty, or errors like
  `Window target pid ... is stale or no longer running`.
- `hermes computer-use doctor` reports `X11 is not reachable`.
- Screenshot analysis says `No LLM provider configured for task=vision`.
- You need to drive a desktop app on a Linux box from an SSH-connected Hermes.

## Architecture (why DISPLAY matters)

Hermes spawns cua-driver as an MCP child process (`cua-driver mcp --no-overlay`)
that **inherits Hermes' own `os.environ`**. On Linux, cua-driver needs X11
credentials in that environment:

- `DISPLAY` (e.g. `:0`) — which X server to talk to
- `XAUTHORITY` (e.g. `$HOME/.Xauthority`) — permission to access it

A bare SSH/TUI session has neither → cua-driver can enumerate windows via
X11 root but fails capture/injection: screenshots come back `0x0`, doctor says
`X11 is not reachable; Set DISPLAY (X11)`.

## Fix: inject DISPLAY via ~/.hermes/.env

Hermes loads `~/.hermes/.env` into `os.environ` at startup (override=True,
`hermes_cli/env_loader.py`), so the display variables must live THERE — not in
the interactive shell's `.bashrc` (child processes spawned by Hermes don't read
shell rc files).

```bash
printf '\n# X11 display for computer_use (cua-driver) over SSH sessions\nDISPLAY=:0\nXAUTHORITY=/home/<user>/.Xauthority\n' >> ~/.hermes/.env
```

**Then restart the Hermes process/session** — the env is captured at process
start; editing `.env` mid-session has no effect on the already-running MCP
child. Verify the running process actually has it:

```bash
tr '\0' '\n' < /proc/<hermes-pid>/environ | grep -E '^DISPLAY|^XAUTHORITY'
```

Check the desktop session first: `loginctl list-sessions` shows which user+seat
owns the X server; `ls ~/.Xauthority` confirms the auth file. Xorg usually runs
on `:0` under lightdm/gdm (`ps aux | grep Xorg`).

## Pitfalls

- **Never start a second cua-driver daemon.** Hermes already runs
  `cua-driver mcp` as its own daemon; running `cua-driver serve` by hand makes
  two daemons fight over the same socket (`~/.cache/cua-driver/cua-driver.sock`)
  → `Window target ... is stale` errors from the Hermes side. Kill the stray one:
  `ps aux | grep 'cua-driver' | grep -v grep` → kill anything that is not
  Hermes' own `mcp` process. Killing Hermes' stale MCP child also works —
  Hermes respawns it on next use.
- **`list_windows` can succeed while `capture` fails** — window enumeration
  needs less X access than pixel capture. Don't conclude the driver works from
  `list_windows` alone; always try a `capture`.
- **The desktop app (`hermes-desktop`) starts WITH DISPLAY** — computer_use
  works there out of the box; it's only SSH/TUI-spawned sessions that lack it.
- Capture with `app=` targeting the frontmost window avoids noisy multi-window
  screenshots.
- **`pkill -f` can kill your own Hermes shell.** The Hermes terminal wrapper
  embeds the command string in its own argv (`bash -c ... eval 'pkill -f
  "FunGen/fungen"' ...`), so `pkill -f <pattern>` matches the shell running it
  and the whole session dies mid-task (exit -9, "killed"). Use exact-name
  matching instead: `pkill -x fungen` (matches the process name only), or
  resolve PIDs first with `pgrep -x`.

## Driving egui / custom-rendered apps (FunGen, Rust/egui UIs)

Rust/egui apps (and some Electron ones) render everything themselves and
expose almost nothing to AT-SPI: `capture mode=som` returns a single
`window` element and zero interactable children, even though the UI is fully
clickable. Don't conclude the app is un-drivable — switch strategy:

- **Background synthetic clicks are ignored by egui.** Use
  `delivery_mode='foreground'` for every click on such apps; background
  `global_input` clicks return `ok` but nothing happens.
- **Drive by vision + coordinates**: take `capture mode=vision`, read the
  layout from `vision_analysis`, then click menu/tab coordinates. Probe
  coordinates iteratively (click → capture → adjust); menu-bar items sit at
  y≈15-25, dropdown items ~25px below their trigger.
- **Keyboard shortcuts from the app's own config are a reliable fallback** —
  egui apps usually have a keymap file (e.g. `prefs.ron` with
  `OpenPreferences: [(key:Comma, command:true, ...)]` → `Ctrl+,`). But note
  synthetic hotkeys may also be ignored until the window has focus; a
  foreground click on the menu first often unblocks them.
- **Native file dialogs pop as a separate window** (`xdg-desktop-portal-gtk`
  via rfd) — `capture app=<app>` still shows the main window; look for the
  portal window in `list_windows` / `xwininfo -root -tree` after triggering
  Open.
- Details + a full worked example: `references/egui-app-driving.md`.

## SIGILL crash on app launch: diagnose instruction-set mismatch

If a GUI app launches fine with no arguments but dies (exit `-4`, no panic
report) when it starts a compute/ML worker, the CPU may not support the
instructions the binary was compiled for:

1. `dmesg | grep -i "trap invalid opcode"` — shows the crashing thread name
   (e.g. `fungen-ml`) and faulting IP. A Rust panic would leave a report;
   kernel SIGILL means raw illegal instruction.
2. `grep -oE "avx2|avx|sse4_2|fma" /proc/cpuinfo | sort -u` — Ivy Bridge-era
   CPUs (i3-3240 etc.) have AVX but **no AVX2**; ONNX Runtime / ML libs
   compiled with AVX2 crash there.
3. `strings <binary> | grep -oE "APP_[A-Z_]+"` — hunt for env-var toggles
   that disable the ML/GPU feature (`*_NO_YOLO`, `*_ML_*`). Often the cleanest
   fix is deleting/blocking the model file so the worker fails gracefully
   (check the app also auto-RE-downloads the model — move the `.part`/`.onnx`
   out of the cache dir to interrupt it).
4. Objdump the faulting offset only if you need proof of which instruction —
   subtract the `in <binary>[base+size]` base from the dmesg IP first;
   virtual address ≠ file offset.

## Screenshot vision needs a vision-capable provider

`capture` returning pixels is only half the battle — interpreting the image
uses the `auxiliary.vision` model (`hermes config get auxiliary`). If it errors
with `No LLM provider configured for task=vision provider=auto`, the main
provider has no vision model available.

- DeepSeek's API (deepseek-v4-flash / deepseek-v4-pro) **rejects `image_url`
  content** (`unknown variant image_url, expected text`) — it is text-only,
  so it can never serve as the vision model.
- Configure a vision-capable provider for `auxiliary.vision`, e.g. OpenRouter
  (has free vision models), Gemini, or Qwen-VL. `hermes setup` walks through
  auxiliary model config.
- Control itself (element tree, clicks, typing) works WITHOUT vision — the
  SOM/AX tree is text; vision only enriches screenshots.

## Verification Sequence

1. `hermes computer-use doctor` — X11 reachability, capture capability.
2. `computer_use list_windows` — window enumeration OK?
3. `computer_use capture mode=vision` — non-zero width/height?
4. `computer_use capture mode=som` — numbered interactable elements appear?
5. Vision interpretation — needs `auxiliary.vision` on a vision-capable
   provider; else expect the "No LLM provider" error on `vision_analysis`.
