Updated 2026-10-11

Extract text from a screenshot: how FoxSnip OCR works

Take a screenshot, then extract text from screenshot pixels that will not select. The model, the reading order, and what never uploads.

By FoxPack

FoxSnip’s plain-text OCR is a local model. You take a screenshot, select the words that will not highlight, and the app lays copyable text back on top of those pixels. Extract text from screenshot is the whole feature. Screenshot to text is the result. Nothing in that pass is uploaded.

What you do

Take a screenshot of a region, a window, or the full display. If the words sit in a picture, a canvas, or someone else’s app, select them and ask for text. FoxSnip recognizes the lines, then opens an overlay on that same selection: a blank page, each line sitting on the box where it was found, ready to drag across and copy.

The overlay is not a second OCR app to learn. It is the screenshot, with the characters made selectable. QR codes in the same image are decoded beside the text. A scrolling screenshot, the full page screenshot, is a different capture: it stitches a long page first. You can run OCR on that tall image afterwards.

The model that ships

Recognition does not call a cloud OCR API, and it does not depend on a system language pack. The build bundles PaddleOCR PP-OCRv5 mobile, exported to ONNX by RapidOCR, under the Apache-2.0 license.

Two networks, both on the CPU:

  1. Detection finds text boxes on the image.
  2. Recognition reads each box into characters.

One recognition model covers simplified Chinese, traditional Chinese, English, and Japanese. The detection weights stay in the official ONNX. The recognition weights are the same model quantized to int8, so detection, recognition, and the character dictionary stay under 20 MB inside the app. The ONNX session is created once and reused. It runs on CPU threads, not on a server.

There is no separate direction classifier in the package. A strip that is taller than it is wide, or a line whose confidence is below 0.4, is rotated 90 degrees clockwise and read again. The better of the two scores is kept. Lines under a 0.15 confidence floor are dropped. Overlapping boxes are suppressed. The survivors are sorted into reading order: top to bottom, and left to right on the same line. An image longer than about 2000 pixels on a side is cut into tiles with 64 pixels of overlap, so a full page screenshot does not have to be one giant tensor. A run that exceeds 12 seconds stops instead of hanging.

Decoding is CTC against the bundled dictionary. Callers receive the string, the box in the original image, and a confidence score. They do not receive raw logits.

QR codes do not use that neural net. They use a classic decoder, so a code does not add model weight.

An earlier Apple Vision path exists in the Mac capture crate and is not connected to the entry the app uses. Shipping OCR is the bundled model, so the same screenshot to text path does not depend on which Vision revision the Mac has loaded.

What stays on the machine

Plain-text OCR stays on the device. It does not spend a Pro credit. Formula recognition, table recognition, chemical-symbol recognition, cloud translation, and share links are the Pro cloud features, and they upload only the selection you send. If you only wanted to extract text from screenshot pixels, you never leave the Mac.

Screenshot Mac, and the other searches

FoxSnip is screen capture software for Mac: a screen capture tool you can install today and use to take a screenshot. A screenshot Windows build and a screenshot PC build are not available yet. Until that installer exists, screenshot Windows and screenshot PC are not jobs this download does.

The jobs it does do on a Mac are the ones above: take a screenshot, make a scrolling screenshot or a full page screenshot, and extract text from screenshot into screenshot to text with local OCR.

Start the 15-day trial

Download the macOS trial. Plus is buy-once. Nothing uploaded by default.

System requirements: macOS 12+ (Apple Silicon & Intel)