what is screen-aware AI?
the short answer
Screen-aware AI is an assistant that can see your screen and use it as context. Instead of copying text out, taking screenshots, and describing what you're looking at, you just ask. The AI already knows what's in front of you.
It's the difference between “I'm on a checkout page for a flight, SFO to Austin, the price says $412 but my card is showing $460, why?” and “why is this charging me more?”
how it works
Every screen-aware assistant does three things:
- Capture. It gets pixels, by taking screenshots, recording the screen, or reading a tab you share.
- Understand. Vision models turn those pixels into structure: text, layout, charts, what changed since last time.
- Answer or act. A language model takes your question plus that screen context and responds. Some tools also watch over time, or take action.
The big fork in the category is where the understanding happens. Most desktop tools upload your screenshots to a cloud model. A smaller set, harness among them, indexes the screen on your own device. Hosted turns can send selected current or recalled evidence, and standard/private modes can send a compact, expiring log of recent changes for continuity. Your screen contains everything: banking, messages, work. Where it gets read is the first question to ask any tool in this category.
what you can do with it
- Ask about what's there. “What does this error mean?” “Summarize this thread.” “Is this a good deal?”
- Ask about what was there. A searchable visual memory means “what was that repo I had open twenty minutes ago?” works.
- Watch for a moment. “Tell me if I get outbid.” “Flag me when this number crosses 2%.” The assistant keeps looking so you don't have to.
- Act on it. Draft the reply, fill in the doc, generate the spreadsheet, using the screen as the source of truth.
the landscape
| tool | runs where | screen reading | notes |
|---|---|---|---|
| harness | any chromium browser, no install | on-device indexing; selected evidence and an optional compact change log leave | watches, remembers, acts; free tier |
| microsoft copilot vision | windows | cloud | sees what you share while sharing is on |
| chatgpt desktop | mac / windows app | cloud | screenshot-based |
| cluely | desktop app | cloud | meeting and interview focus |
| highlight ai | desktop app | cloud | context from screen plus audio |
| precogni | desktop app | snapshot at ask-time | privacy-positioned |
| ai cowork | desktop, open source | local | free, DIY |
For the head-to-heads, see the linked comparisons above, or the full 2026 roundup.
why the browser
Desktop screen recorders see everything, always, which is exactly why most people never turn them on. A browser tab you explicitly share is a permission model people already understand from every video call they've ever been on. No install, no always-on recorder. Share what you choose, stop when you want.
That's the version harness builds: screen-aware AI in your browser, free to try, with screen capture and indexing that stay on your device.
common questions
what is screen-aware AI?
screen-aware AI is an assistant that can see your screen and use it as context, so you stop retyping and describing what’s in front of you. harness is the browser-native take: share a tab, a window, or your whole screen, and it answers questions about what’s actually there, watches for the moment you care about, and acts when you ask. capture and indexing run on your device. hosted turns can send selected current or recalled evidence; standard and private modes can also send a compact, expiring log of recent changes.
is there an AI that can see my screen?
yes. harness is an AI that can see your screen: open it in any chromium browser, share a tab, a window, or your whole screen, and ask about what’s in front of you. no install, no screenshots to feed it. it can also keep watching after you look away and flag you the moment something you care about happens. screen indexing runs on your device; hosted turns can send selected evidence and, outside confidential mode, a compact expiring change log.
is harness actually private, or is that marketing?
harness is built for privacy, and how private is your call. your full screen never streams up, and your screen history and memory databases stay on your device. capture, indexing, and retrieval run in your browser. standard and private hosted modes can send selected evidence plus a compact, expiring text/audio log of recent changes. confidential mode keeps that log local and uses a TEE route. a local endpoint keeps model execution on your hardware, though hosted orchestration still transits harness.
which AI is doing the answering?
two layers. the watching layer is small vision models running in your browser (a CLIP encoder plus a compact captioner), and it indexes your screen locally. the answering layer is your call: run a model locally, connect an OpenAI-compatible endpoint, bring a venice or bankr key, or use managed credit routed to frontier models (gemini, claude, gpt, grok), no key needed.
what does “bring your own key” mean?
harness reaches AI through Venice and Bankr, or through an OpenAI-compatible endpoint you configure. if you have a venice or bankr key, plug it in and harness spends your own credit instead of ours. the open tier is free this way, and fully free if you run a model locally on your own machine. prefer not to bring a key? use managed credit instead, pay-as-you-go, starting with $5 free. resident adds encrypted backup of your memory and conversations, restorable on another device.
can it look back at a screen from earlier, or watch something for me?
yes, both. harness keeps a short, searchable visual memory of what it has seen, so you can ask about a screen from ten minutes ago without having it open again. you can also point it at something and say watch this for twenty minutes, note when the number crosses a threshold, then write me a summary. it keeps an eye on it and reports back. it never clicks or types for you, it watches and tells you what it saw.
do i need to install anything?
no. harness runs in any modern chromium browser (chrome, brave, arc, edge). open the tab, share what you want it to see, ask. that’s the whole install.
what does it remember about me?
your preferences (how you like answers shaped), recurring contexts (you trade on hyperliquid, you study for the bar), and the thread of your work: what you asked yesterday, what you were on. not your screen, not screenshots, not your typing. just the gist, stored in a database on your own device, wipe-able at any time.
is it open source? can i run it myself?
the harness client is being open-sourced. you’ll be able to read every line that watches your screen and holds your memory, and run the whole thing on your own infrastructure if you’d rather not use the hosted version. the hosted product and the self-hosted one are the same client. that’s the point: you don’t have to trust us, you can check.
what if i cancel?
you drop back to the open tier and keep everything that runs on your device: your local memory, your own model or key, the floating window, and access to metered managed inference. what pauses is the Resident layer, including encrypted cloud backup and other paid capabilities. your memory is local, so it’s never held hostage. for a full delete, one click in settings.