Is CroppickVoice really free?
Yes. Unlimited dictation on your device, the full on-device model catalog, voice modes and core Crop & Speak are free — no word caps, no account, no trial clock. A one-time €39 Pro upgrade adds power features on top; nothing that is free today ever moves behind the paywall.
Does my voice or screen ever leave my computer?
Not unless you deliberately send it to a cloud service you chose. By default speech becomes text on your device, screenshots stay in a local folder, and the app tracks nothing about you. The network is used only to download the models you pick (once each), for an optional signed update check, for a one-time activation if you buy Pro — and by the opt-in features: text polish, which can stay fully local via the built-in model or your own Ollama server, Pro’s optional cloud transcription through your own account, any Pro output action you point at a webhook of your choosing, and whatever you let a connected AI app read through Pro’s agent bridge — that content joins your conversation with the app’s own AI. Everything cloud is off until you set it up.
Does it work offline?
Yes. After a one-time model download, dictation and Crop & Speak run entirely on your computer — pull the network cable and they keep working. The network is touched only for things you choose: downloading a model you picked, the optional signed update check, Pro’s one-time activation, and any cloud AI you deliberately turn on.
What is Crop & Speak?
Press Alt+`, drag over any region of your screen, mark it up with arrows, boxes or a pen while you talk, and CroppickVoice delivers the annotated image plus your transcribed words into the focused app — image and text for chats, docs and editors, text and file path for terminals. Saying it and showing it becomes one gesture instead of two apps.
Can I dictate into ChatGPT or Claude?
Yes — it’s one of the most common uses. Put your cursor in the chat box, press Ctrl+`, and talk: your words appear right there, ready to edit before you send. And when the question is about something on your screen, Crop & Speak sends the annotated screenshot and your spoken words together.
Can I use it just for screenshots, without voice?
Yes. Alt+1 captures without recording: drag over any part of the screen, mark it up with arrows, boxes or a pen, and it lands wherever you paste — no microphone involved.
Why did I sometimes get a file path instead of the image?
Because Crop & Speak adapts to the app you paste into, and apps differ in what they can accept. Apps that understand pictures — chats, documents, email, editors — receive your words and the annotated image together. Text-only places like terminals can’t show a picture, so they get your words plus the image’s file location, which is usually what you want there. Got a path where you wanted the picture? Hover the button, open History, and click the entry — that second delivery often lands the image + text in apps where the first paste couldn’t. You can also set a fixed preference in Settings → Crop & Speak → Delivery.
Which languages does it support?
The built-in model catalog covers the majority of languages across Europe, Asia and the Americas. Pick a model in Settings — lighter ones for speed on modest hardware, larger ones for maximum accuracy — and your language is auto-detected, or pin it explicitly.
Can I see the words while I’m still speaking?
Yes — turn on live captions. Free shows your words at each natural pause. Pro adds a streaming cadence that paints the caption word by word while you talk — fully on-device where a streaming model covers your language, or through your own cloud speech account. Either way the finished text lands in your app only when you stop, so nothing half-revised is ever typed.
Can a dictation land somewhere besides the cursor — a file, Slack, a script?
With Pro, yes. Output actions send each finished dictation somewhere in addition to (or instead of) the normal paste: append it to a file — a journal that grows as you speak — post it to a webhook (Slack, Discord, ntfy and Home Assistant presets are built in), or pipe it into a local command. Actions are routed per voice mode, everything is off until you set it up, and webhook tokens are sealed in the same encrypted local vault as your account keys.
Can AI apps like Claude Code or Cursor pull things from CroppickVoice?
With Pro, yes. Turn on the agent bridge and an AI app you already use — Claude Code, Claude Desktop, Cursor, VS Code — can connect to CroppickVoice, which acts as a local MCP server on your computer. (MCP, the Model Context Protocol, is the open standard that lets AI apps use tools offered by programs on your machine; setup is one click from Settings.) A connected app can read your latest Crop & Speak capture, search your dictation history, or ask you for a fresh snip — that request opens the normal capture overlay, nothing is shared until you finish it, you can redact first, and Esc declines. The bridge is off by default, every ability has its own switch, and it only answers on your own computer — it never connects out. Worth knowing: whatever a connected app reads becomes part of your conversation with that app’s AI, which usually runs in the cloud.
Is this text to speech?
The other way around. CroppickVoice is speech to text — it types what you say. It doesn’t read text out loud.
Windows says "Windows protected your PC" — why?
The installer is new and doesn’t carry its Windows publisher certificate yet, so SmartScreen plays it safe. Click "More info" → "Run anyway". Updates themselves are digitally signed and verified before they install, and every download’s checksum is published so you can verify what you got.
When do macOS and Linux ship?
The app is built cross-platform and both ports already run from source; polished packaged builds are being validated. Windows 10/11 is ready today — the macOS and Linux builds will appear on the download page the moment they’re ready.
Is Pro a subscription?
No. Pro is €39 once — buy it once and every 1.x update is included. Your key activates with one quick online check, then works fully offline — no account, no subscription, never a recurring check.
Is there a team or enterprise plan?
Coming soon — team licensing with seats bought together and one invoice for the company is in the works, and pricing hasn’t been announced yet. Until then it stays simple: each person gets their own €39 Pro license, and the free tier is free for commercial use, so a whole team can start today.
Do I need a powerful computer?
No. The standard model keeps up with your speech on an ordinary computer. If you do have a graphics card — NVIDIA, AMD, Intel or Apple — the app uses it for extra speed.
Why is it using a lot of memory (RAM)?
That’s the price of working offline: the speech model is held in memory while the app runs, which is exactly what makes transcription instant and private. How much depends on your choice in Settings — light models keep the footprint small, the big high-accuracy ones take considerably more room. And if you turn on the built-in local AI polish, that second model loads into memory too — running both at once is when usage looks large. Want it leaner? Pick a lighter speech model, or run polish through your own Ollama server instead of the built-in one. Memory is released when you quit the app.