How-to

A macOS OCR keyboard shortcut: extract text from any screenshot

macOS has no built-in keyboard shortcut that OCRs a screen region straight to your clipboard. Live Text handles images you can open in Photos or Preview, but not a video frame, a screen-share, or an app that blocks selection. Khint adds the missing gesture: press Cmd+Shift+K, pick Extract text, drag a box, and the recognized text is on your clipboard. The capture is taken by Apple's own screenshot tool, so the app needs Accessibility permission and never Screen Recording.

Some text refuses to be selected: a screenshot a colleague pasted into Slack, a frame paused in a recorded demo, an error dialog with no copy button, a scanned invoice, a slide in someone else's screen-share. You can read every word. You just can't take any of them. The usual workaround is to retype it, and the usual result is a typo in an error code you were about to search for.

OCR turns those pixels back into text. The interesting question is not whether macOS can do it, but how few gestures it takes to get the words onto your clipboard. Here is the setup that makes it one: shortcut, drag, paste, with Khint.

The gesture, end to end

  1. Press Cmd+Shift+K

    The palette opens over whatever you were looking at. You don't switch apps, and the window behind you keeps focus, so nothing you had selected is lost.

  2. Pick Extract text

    Under Capture. This is the row that opens no window and produces no file: it exists to hand you words.

  3. Drag over the text

    The macOS region selector takes over, the same crosshair as Cmd+Shift+4. Draw a box, release.

  4. It is already on your clipboard

    A small heads-up display confirms the copy and shows the first line of what it read. Nothing opens, nothing waits for you. Paste it wherever you were going.

Why there is no dedicated OCR hotkey

Khint used to bind Cmd+Shift+O to capture. That binding was removed in April 2026, on purpose, and it is worth saying why rather than pretending capture never had its own key.

A second global hotkey buys you one saved keystroke and costs you a permanent conflict surface: every app you install is competing for the same four-key combinations, and a shortcut that silently stops working is worse than one that never existed. Capture is now a row in the palette you already opened, which means it also sits next to your AI actions, your workflows, and your integrations. The capture and what you do with the result are one flow instead of two.

How this differs from macOS Live Text

Live Text is genuinely good where it applies. Open an image in Photos, Preview, or Quick Look, hover, and you can usually select the text inside it. If that covers your case, you do not need anything else.

The gaps are the reason a shortcut exists. Live Text needs an image, in an app that supports it, that you can open. It does not help you with a paused video frame, a colleague's screen-share, a Figma export, an Electron app that draws its own text, or a dialog that will not stay open while you go find it in Preview. A global shortcut works identically in all of them, because it operates on the screen rather than on a file. And it copies straight to the clipboard rather than asking you to select and copy by hand.

When you want more than the raw text

Sometimes the words are not the point. The palette's Capture & askrow takes the same region and opens a short conversation about it: "summarize this table", "what is this stack trace actually complaining about", "turn this slide into bullets I can paste into the deck". The image stays pinned for the whole exchange, so follow-up questions do not need a second capture.

Because both rows live in the same palette as your agents, the result is one step from a transformation rather than a dead end. Extract a block of handwritten notes from a whiteboard photo, then turn it into a Jira ticket without the text ever passing through your keyboard.

Things worth capturing

  • An error message in a dialog with no copy button, before you search for it and mistype the code.
  • A command or code snippet shown in a video tutorial or a conference talk.
  • A table locked inside a scanned PDF, where every cell would otherwise be retyped.
  • A slide in a screen-share you have no file for and will not get one of.
  • A whiteboard photo from a workshop, on the way to becoming requirements.
  • A receipt, a serial number, or a label you photographed instead of writing down.

What it costs, and what it refuses to do

A capture costs roughly one credit, out of 300 a month on the free tier. Cancelling the region select costs nothing at all: back out of the crosshair and no call is made, because the capture happens before anything is billed. Only one capture runs at a time, so a second shortcut press while a region is open will not start a competing flow.

Two honest limits. Handwriting is recognized unevenly, and a photograph of cursive on a whiteboard will need a human pass. And the region is what you drew: if half a sentence sits outside the box, half a sentence is what comes back. Drawing generously costs nothing.

Common questions

What's the keyboard shortcut for OCR on a Mac?

macOS has no built-in one. In Khint, OCR lives in the palette: press Cmd+Shift+K, then choose 'Extract text'. There is deliberately no separate always-on hotkey, so capture and your AI actions share one shortcut instead of competing for two.

Does it need Screen Recording permission?

No. The region is captured by screencaptureui, Apple's own screenshot tool, which runs with your permissions rather than the app's. Khint asks for Accessibility permission, which it needs to paste results back at your cursor, and never asks for Screen Recording. It has no background capture of any kind.

Does it work on images, PDFs, and video frames?

Yes. It reads whatever is on your screen inside the region you drag: an image, a paused video frame, a scanned PDF, a slide in a screen-share, or an app that blocks text selection. You can also point it at an image file directly instead of capturing a region.

Is my screenshot uploaded or stored?

The region is captured on your Mac. That single image is sent once to read the text and is not stored after the call returns. There is no background screen recording and no capture history.

How is this different from Apple's Live Text?

Live Text works on an image you can open in a supporting app, such as Photos or Preview. A capture shortcut works on the screen itself, so it also covers video frames, screen-shares, and apps that draw their own text. It also copies straight to the clipboard instead of requiring a manual select-and-copy.

How many captures can I do for free?

The free tier includes 300 credits a month and a capture costs about 1 credit. Paid plans give you a bigger monthly balance: 900 on Lite, 4,000 on Pro, 7,000 on Max. Cancelling a region select costs nothing.

Try it in your own workflow

Khint runs your prompts on selected text in any Mac app, from one shortcut. Free with 300 credits a month, about 10 AI actions a day.