Talk to type.
Snip to show.

Free speech to text and screenshot annotation — your voice types, your screen talks, nothing leaves the machine.

free · no account · works offline · Windows 10/11

01 · voice

Speech to text in any app — offline dictation in most languages

02 · crop & speak

Show it. Say it. Send it.

how capture & delivery work →

03 · yours

Dress it. Move it. Hide it.

theme
layout
tray

Grab the button and drag it anywhere — the real one moves the same way and remembers its spot on every screen. When idle it politely fades, and wakes as your mouse comes near — try it.

all the knobs →

04 · private

Unplug the internet.

what touches the network, exactly →

That’s the whole setup.

questions? the FAQ →

voice — the details

Speak at ~150 words a minute instead of typing at 40. CroppickVoice listens on one keyboard shortcut and types wherever your cursor already is.

One shortcut, any app

Press Ctrl+` in your email, browser, document or chat — the words land at your cursor, already on the clipboard as backup.

Tap or hold to talk

Toggle, push-to-talk, or a hybrid that decides from how long you hold. Optional live captions show your words as you pause — and Pro can stream them word by word while you speak.

Voice modes that rewrite

Polish fixes grammar. Prompt turns rambling into a clear AI request. Add your own modes, each with its own optional shortcut — free on the built-in local model or your own Ollama server; Pro connects your OpenAI / Claude account.

Most languages, on-device

A selection of speech models covering the majority of languages across Europe, Asia and the Americas — all running locally. Auto-detect or pin a language.

Clipboard-safe paste

Auto-insert borrows your clipboard and puts the previous content back. Nothing you copied gets lost.

History at a hover

Hover the button for recent transcripts and shots — search, filter, re-paste, polish, or drag an image straight into another app.

the model catalog — all free, all local
  • Most languages covered a growing selection of models across Europe, Asia and the Americas
  • Light to heavyweight quick ones for any computer, larger ones for maximum accuracy
  • Swap anytime pick in Settings — models download on demand, then run offline forever

crop & speak — the details

Words alone take paragraphs; a bare screenshot still needs explaining. Crop & Speak does both in one motion — and skips the screenshot → save → attach → type-it-all-up dance.

  1. 1

    Press Alt `

    The screen dims on every monitor and recording starts. ( Alt 1 captures without voice.)

  2. 2

    Drag, mark up, talk

    Arrow, pen, box, highlighter — circle the problem while you describe it. Undo included.

  3. 3

    It lands where you paste

    Chats, docs and editors get your words + the annotated image. Terminals get text + the image’s file path. One paste, the right format.

people use it to
  • ask AI about an error
  • report a bug
  • give design feedback
  • file a support ticket
  • query an invoice line
  • walk a colleague through a form

make it yours — the details

It floats above everything, so it behaves like a good guest: theme it, scale it, fade it, drag it anywhere, or tuck it into the tray and do everything from the keyboard.

Shortcuts, your way

Dictation, Crop & Speak and capture-only each get their own keyboard shortcut — change any of them to whatever your fingers prefer.

At home on every screen

The button remembers its spot on each monitor, flips to a vertical layout if you prefer edges, and returns to a visible place if its screen ever disappears.

Works over admin windows

Your shortcuts keep working over programs run as administrator — and the app tells you when pasting there needs an extra step.

Small install, models on demand

Speech models download once with visible progress, then stay on your machine. Installer or unzip-and-run — your pick.

privacy — the details

By default nothing you say, show or type leaves your computer. Don’t take our word for it — pull the network cable: dictation and Crop & Speak keep working.

Always local

Your voice · your screenshots · your history · the built-in speech-to-text engine

Download-only

The speech & polish models, fetched once when you pick them · a tiny signed version check for updates (can be turned off)

Only if you turn it on

AI polish through your own OpenAI / Claude account sends that text; Pro’s optional cloud transcription sends audio to the speech provider you configure — clearly labelled, off until you enable them. Prefer local? The built-in model or your own Ollama server keeps polish offline too. Pro’s output actions and its agent bridge for AI apps are the same story: nothing moves unless you set them up, and only where you point them.

Safe, quiet updates

The app checks a small, digitally signed version file — the signature is verified before anything is trusted, every download is checksum-verified, and if anything doesn’t match, the update is simply rejected. You can turn checks off, or block updates entirely.

No tracking. No account. Nothing phoning home — from the app or from this site. Read the full network-use table and the free-forever promise →

pricing — free vs pro

Everything that’s free today stays free — that’s a promise, not a trial.

Free

€0 forever

  • Unlimited on-device dictation — no word caps
  • The full on-device model catalog — Europe, Asia & the Americas covered
  • Core Crop & Speak — capture, annotate, narrate, smart delivery
  • Voice modes: Polish, Prompt + your own
  • AI polish — built-in local model or your own Ollama server
  • Recent history — search, filter, drag out
  • Push-to-talk, live captions, safe updates
Download for Windows

v1.2.0 ·all downloads & checksums

Pro

€39 once · every 1.x update included

  • Agent bridge for AI apps (MCP) — let Claude Code, Claude Desktop, Cursor or VS Code read your latest crop, search your dictation history, or ask you for a snip you approve — answered on your computer only
  • Streaming captions — watch words appear while you’re still speaking, on-device or through your own cloud speech account
  • Output actions — send each dictation to a file, a webhook (Slack, Discord, ntfy, Home Assistant…) or a script, routed per voice mode
  • Cloud transcription, your keys — send audio to your own cloud speech provider when you want maximum accuracy
  • Cloud polish, your keys — run text enhancement through your own OpenAI or Claude account
  • A different AI per mode — each voice mode can use its own model or provider
  • Per-app behavior — different settings per program, down to the keystroke sent after a paste
  • History that survives restarts — keep it on disk, from 50 entries up to unlimited
  • Your own cleanup rules — teach the transcript filter your phrases to strip or rewrite
Get Pro

Activate once, then it works fully offline — no account · price + VAT where applicable · 30-day refund

Teams & Enterprise

Coming soon pricing to be announced

  • Pro for the whole team — seats bought together
  • One invoice for the company, keys delivered in one go
  • The free tier is free at work — a team can start today

Checkout opens when pricing is announced · until then each person gets their own Pro license

Still local-first: by default your words and screenshots never leave your device. Out of the box Pro’s only network touch is a one-time license activation (your key plus an anonymous device code, never your content); the cloud AI options are opt-in, clearly labelled, and use your own accounts. The free-forever promise lists what can never move behind the paywall.

fair questions

Is CroppickVoice really free?

Yes. Unlimited dictation on your device, the full on-device model catalog, voice modes and core Crop & Speak are free — no word caps, no account, no trial clock. A one-time €39 Pro upgrade adds power features on top; nothing that is free today ever moves behind the paywall.

Does my voice or screen ever leave my computer?

Not unless you deliberately send it to a cloud service you chose. By default speech becomes text on your device, screenshots stay in a local folder, and the app tracks nothing about you. The network is used only to download the models you pick (once each), for an optional signed update check, for a one-time activation if you buy Pro — and by the opt-in features: text polish, which can stay fully local via the built-in model or your own Ollama server, Pro’s optional cloud transcription through your own account, any Pro output action you point at a webhook of your choosing, and whatever you let a connected AI app read through Pro’s agent bridge — that content joins your conversation with the app’s own AI. Everything cloud is off until you set it up.

Does it work offline?

Yes. After a one-time model download, dictation and Crop & Speak run entirely on your computer — pull the network cable and they keep working. The network is touched only for things you choose: downloading a model you picked, the optional signed update check, Pro’s one-time activation, and any cloud AI you deliberately turn on.

What is Crop & Speak?

Press Alt+`, drag over any region of your screen, mark it up with arrows, boxes or a pen while you talk, and CroppickVoice delivers the annotated image plus your transcribed words into the focused app — image and text for chats, docs and editors, text and file path for terminals. Saying it and showing it becomes one gesture instead of two apps.

Can I dictate into ChatGPT or Claude?

Yes — it’s one of the most common uses. Put your cursor in the chat box, press Ctrl+`, and talk: your words appear right there, ready to edit before you send. And when the question is about something on your screen, Crop & Speak sends the annotated screenshot and your spoken words together.

Can I use it just for screenshots, without voice?

Yes. Alt+1 captures without recording: drag over any part of the screen, mark it up with arrows, boxes or a pen, and it lands wherever you paste — no microphone involved.

Why did I sometimes get a file path instead of the image?

Because Crop & Speak adapts to the app you paste into, and apps differ in what they can accept. Apps that understand pictures — chats, documents, email, editors — receive your words and the annotated image together. Text-only places like terminals can’t show a picture, so they get your words plus the image’s file location, which is usually what you want there. Got a path where you wanted the picture? Hover the button, open History, and click the entry — that second delivery often lands the image + text in apps where the first paste couldn’t. You can also set a fixed preference in Settings → Crop & Speak → Delivery.

Which languages does it support?

The built-in model catalog covers the majority of languages across Europe, Asia and the Americas. Pick a model in Settings — lighter ones for speed on modest hardware, larger ones for maximum accuracy — and your language is auto-detected, or pin it explicitly.

Can I see the words while I’m still speaking?

Yes — turn on live captions. Free shows your words at each natural pause. Pro adds a streaming cadence that paints the caption word by word while you talk — fully on-device where a streaming model covers your language, or through your own cloud speech account. Either way the finished text lands in your app only when you stop, so nothing half-revised is ever typed.

Can a dictation land somewhere besides the cursor — a file, Slack, a script?

With Pro, yes. Output actions send each finished dictation somewhere in addition to (or instead of) the normal paste: append it to a file — a journal that grows as you speak — post it to a webhook (Slack, Discord, ntfy and Home Assistant presets are built in), or pipe it into a local command. Actions are routed per voice mode, everything is off until you set it up, and webhook tokens are sealed in the same encrypted local vault as your account keys.

Can AI apps like Claude Code or Cursor pull things from CroppickVoice?

With Pro, yes. Turn on the agent bridge and an AI app you already use — Claude Code, Claude Desktop, Cursor, VS Code — can connect to CroppickVoice, which acts as a local MCP server on your computer. (MCP, the Model Context Protocol, is the open standard that lets AI apps use tools offered by programs on your machine; setup is one click from Settings.) A connected app can read your latest Crop & Speak capture, search your dictation history, or ask you for a fresh snip — that request opens the normal capture overlay, nothing is shared until you finish it, you can redact first, and Esc declines. The bridge is off by default, every ability has its own switch, and it only answers on your own computer — it never connects out. Worth knowing: whatever a connected app reads becomes part of your conversation with that app’s AI, which usually runs in the cloud.

Is this text to speech?

The other way around. CroppickVoice is speech to text — it types what you say. It doesn’t read text out loud.

Windows says "Windows protected your PC" — why?

The installer is new and doesn’t carry its Windows publisher certificate yet, so SmartScreen plays it safe. Click "More info" → "Run anyway". Updates themselves are digitally signed and verified before they install, and every download’s checksum is published so you can verify what you got.

When do macOS and Linux ship?

The app is built cross-platform and both ports already run from source; polished packaged builds are being validated. Windows 10/11 is ready today — the macOS and Linux builds will appear on the download page the moment they’re ready.

Is Pro a subscription?

No. Pro is €39 once — buy it once and every 1.x update is included. Your key activates with one quick online check, then works fully offline — no account, no subscription, never a recurring check.

Is there a team or enterprise plan?

Coming soon — team licensing with seats bought together and one invoice for the company is in the works, and pricing hasn’t been announced yet. Until then it stays simple: each person gets their own €39 Pro license, and the free tier is free for commercial use, so a whole team can start today.

Do I need a powerful computer?

No. The standard model keeps up with your speech on an ordinary computer. If you do have a graphics card — NVIDIA, AMD, Intel or Apple — the app uses it for extra speed.

Why is it using a lot of memory (RAM)?

That’s the price of working offline: the speech model is held in memory while the app runs, which is exactly what makes transcription instant and private. How much depends on your choice in Settings — light models keep the footprint small, the big high-accuracy ones take considerably more room. And if you turn on the built-in local AI polish, that second model loads into memory too — running both at once is when usage looks large. Want it leaner? Pick a lighter speech model, or run polish through your own Ollama server instead of the built-in one. Memory is released when you quit the app.