Clicky is a screen-aware voice buddy for the Mac: press a hotkey, talk, and it reads a screenshot of your screen to answer out loud or run a task across your apps. OpenCues removes the voice round-trip and the screenshot. You type _ inside the text you are already writing and the answer lands right there, computed from the text itself rather than from a picture of your screen.
The screenshot is the difference
Both tools are hotkey-first and want to help without you leaving the app you are in. The difference is what they read and where the result goes. Clicky captures your screen and listens to your voice, then answers by speaking, drawing an arrow at the button you need, or dispatching an agent to go do the task. OpenCues reads the text field you are inside and writes back into it: 5 miles in km _ becomes 8 km in the sentence you were writing, with nothing spoken, screenshotted, or pasted.
In the 3-axis terms this site uses: a voice answer or an agent run has real end-lag (listen, read it back, or wait for the task, then get what you need into the text yourself), while an inline fill has close to none, because the destination was declared by where you put the _.
Where Clicky is genuinely stronger
- It sees everything. Screen capture plus accessibility means Clicky works over any visible app with no host to integrate, including native apps OpenCues has no adapter for. OpenCues works inside its supported hosts.
- Voice and tutoring. Clicky is ambient and conversational, and it can point an animated cursor at the exact control you need in an unfamiliar app like DaVinci Resolve or Figma. OpenCues is silent and typed; it does not talk or draw on screen.
- Agentic tasks across apps. Say the word and Clicky will go build a small app, draft an email, file a Linear ticket, or check your calendar. OpenCues agent tasks operate on your text buffer, not by driving other applications.
Where OpenCues goes further
- It lands in your text. The answer appears where you were typing, with no voice to transcribe and no paste step. For the tenth lookup while drafting, that compounds.
- Nothing leaves as a picture. OpenCues computes on the text you typed and never captures your screen; Clicky sends a screenshot to its cloud on every hotkey press. OpenCues also tokenises identifiers so PII stays off cloud providers.
- Your model, your key, open source. OpenCues is Apache-2.0 and model agnostic (Cerebras, Groq, OpenAI, Anthropic, Gemini, OpenRouter, local Ollama). Clicky runs a fixed hosted pipeline behind its subscription; the original Clicky repo is MIT, but the shipping product is closed.
- The passive layer. Cues dim words with alternatives as you compose and grammar correction runs continuously, and any inline edit is revertable per word; a summoned voice buddy has no equivalent, since it only acts when you call it.
The same task, both ways
Concretely: mid-email, you need a conversion. The Clicky route is hold the hotkey, say "what is 5 miles in kilometres", hear "about 8 kilometres", and type it in yourself. The OpenCues route is typing 5 miles in km _ where the number belongs and continuing to type while it resolves; if it cannot resolve, the _ just stays visible and harmless. Clicky is the better tool when the thing you want is not text at all, for instance being shown where a menu is, or when you want an agent to go operate three apps for you. OpenCues wins whenever the destination is the field you are already in. That is why the two coexist: a screen buddy for showing and doing, OpenCues for the text you are inside.
Where Clicky struggles
The screenshot is the tax: every answer starts from a picture of your screen sent to Clicky's cloud, and voice output still has to be typed in by you before it counts as text. There is no bring-your-own-key or local option in the shipping product, so the model and the pipe are both Clicky's, and it is macOS only today with Windows on a waitlist.
By the numbers
Clicky adds one voice channel and one agent surface but no inline text edit: a spoken answer is zero new windows yet a full round-trip back to your caret, and nothing surfaces unprompted in the field you are writing. OpenCues measures one step to an answer with zero end-lag inline. Counts and sequences in the full comparison.
Related: OpenCues vs Raycast AI · OpenCues vs Windows Click to Do · AI in any text field · What is a blank?