harness

what is screen-aware AI?

Updated August 2026

the short answer

Screen-aware AI is an assistant that can see your screen and use it as context. Instead of copying text out, taking screenshots, and describing what you're looking at, you just ask. The AI already knows what's in front of you.

It's the difference between “I'm on a checkout page for a flight, SFO to Austin, the price says $412 but my card is showing $460, why?” and “why is this charging me more?”

how it works

Every screen-aware assistant does three things:

  • Capture. It gets pixels, by taking screenshots, recording the screen, or reading a tab you share.
  • Understand. Vision models turn those pixels into structure: text, layout, charts, what changed since last time.
  • Answer or act. A language model takes your question plus that screen context and responds. Some tools also watch over time, or take action.

The big fork in the category is where the understanding happens. Most desktop tools upload your screenshots to a cloud model. A smaller set, harness among them, reads the screen on your own device and sends only a focused slice of the relevant region when you actually ask something. Your screen contains everything: banking, messages, work. Where it gets read is the first question to ask any tool in this category.

what you can do with it

  • Ask about what's there.“What does this error mean?” “Summarize this thread.” “Is this a good deal?”
  • Ask about what was there.A searchable visual memory means “what was that repo I had open twenty minutes ago?” works.
  • Watch for a moment.“Tell me if I get outbid.” “Flag me when this number crosses 2%.” The assistant keeps looking so you don't have to.
  • Act on it. Draft the reply, fill in the doc, generate the spreadsheet, using the screen as the source of truth.

the landscape

toolruns wherescreen readingnotes
harnessany chromium browser, no installon-device; only a focused slice leaveswatches, remembers, acts; free tier
microsoft copilot visionwindowscloudsees what you share while sharing is on
chatgpt desktopmac / windows appcloudscreenshot-based
cluelydesktop appcloudmeeting and interview focus
highlight aidesktop appcloudcontext from screen plus audio
precognidesktop appsnapshot at ask-timeprivacy-positioned
ai coworkdesktop, open sourcelocalfree, DIY

For the head-to-heads, see the linked comparisons above, or the full 2026 roundup.

why the browser

Desktop screen recorders see everything, always, which is exactly why most people never turn them on. A browser tab you explicitly share is a permission model people already understand from every video call they've ever been on. No install, no always-on recorder. Share what you choose, stop when you want.

That's the version harness builds: screen-aware AI in your browser, free to try, with screen reading that stays on your device.

common questions

what is screen-aware AI?

screen-aware AI is an assistant that can see your screen and use it as context, so you stop retyping and describing what’s in front of you. harness is the browser-native take: share a tab, a window, or your whole screen, and it answers questions about what’s actually there, watches for the moment you care about, and acts when you ask. unlike desktop tools that upload screenshots, harness reads your screen on your own device and sends only a focused slice of the relevant part when you ask a question.

is there an AI that can see my screen?

yes. harness is an AI that can see your screen: open it in any chromium browser, share a tab, a window, or your whole screen, and ask about what’s in front of you. no install, no screenshots to feed it. it can also keep watching after you look away and flag you the moment something you care about happens. the models that read your screen run on your own device, so the pixels stay with you.

is harness actually private, or is that marketing?

harness is built for privacy, and how private is your call. your full screen never streams up, your memory never leaves your device, and your full conversation history is saved there too. the AI that decides what matters runs in your browser, on your machine, and trims what goes out. when you ask a question, what travels is your prompt plus a focused slice of the relevant part of your screen, nothing more. by default that goes through Venice on an anonymized route, so the call can’t be tied back to you. want the content private too, pick a TEE model, or run a model locally so nothing leaves at all. and soon the client is open source, so you won’t have to take our word for any of this.

which AI is doing the answering?

two layers. the watching layer is small vision models running in your browser (a CLIP encoder plus a compact captioner), they read your screen locally and never leave your machine. the answering layer is your call: run a model locally on your own hardware, bring a venice or bankr key, or use managed credit and we route to frontier models (gemini, claude, gpt, grok) through Venice’s anonymized infrastructure, no key needed.

what does “bring your own key” mean?

harness reaches AI through Venice and Bankr, the routers that cover both open models and the closed frontier ones. if you have a venice or bankr key, plug it in and harness spends your own credit instead of ours. the open tier is free this way, and fully free if you run a model locally on your own machine. prefer not to bring a key? use managed credit instead, pay-as-you-go, starting with $5 free. resident adds encrypted backup of your memory and conversations, restorable on any device.

can it look back at a screen from earlier, or watch something for me?

yes, both. harness keeps a short, searchable visual memory of what it has seen, so you can ask about a screen from ten minutes ago without having it open again. you can also point it at something and say watch this for twenty minutes, note when the number crosses a threshold, then write me a summary. it keeps an eye on it and reports back. it never clicks or types for you, it watches and tells you what it saw.

do i need to install anything?

no. harness runs in any modern chromium browser (chrome, brave, arc, edge). open the tab, share what you want it to see, ask. that’s the whole install.

what does it remember about me?

your preferences (how you like answers shaped), recurring contexts (you trade on hyperliquid, you study for the bar), and the thread of your work: what you asked yesterday, what you were on. not your screen, not screenshots, not your typing. just the gist, stored in a database on your own device, wipe-able at any time.

is it open source? can i run it myself?

the harness client is being open-sourced. you’ll be able to read every line that watches your screen and holds your memory, and run the whole thing on your own infrastructure if you’d rather not use the hosted version. the hosted product and the self-hosted one are the same client. that’s the point: you don’t have to trust us, you can check.

what if i cancel?

you drop back to the open tier and keep everything that runs on your device: your local memory, your own model or key, the floating window. what pauses is the paid layer, the cloud backup, or managed AI. your memory is local, so it’s never held hostage. for a full delete, one click in settings.

try harness