Speech-to-Markdown

FreeStarting price
0Popularity
Speech-to-Markdown featured image

About Speech-to-Markdown

Speech-to-Markdown is a macOS menu-bar application and an iPhone/iPad app that converts voice input into structured markdown documents entirely offline. The macOS app uses whisper.cpp for on-device speech recognition and a local LLM server to format the transcribed text into clean markdown, plain text, or HTML in real time. Users can choose between two modes: Global Dictation, which types transcribed speech directly into any application via a hotkey, and Agent Mode, which streams voice input into a structured document that updates continuously as the user speaks. The app supports session history, allowing users to restore previous documents and continue recording. Output formats can be toggled between Markdown, plain text, and HTML, with each format applying specific formatting rules. The macOS app requires a local LLM server running on the user’s device, supporting multiple open-source LLM servers such as omlx, Ollama, LM Studio, and llama.cpp. The iOS app operates entirely offline using Apple Intelligence and does not require additional dependencies.

Key features

  • On-device speech recognition with whisper.cpp
  • Real-time streaming transcription and formatting
  • Global Dictation mode for typing directly into any app
  • Agent Mode for structured document creation
  • Support for Markdown, plain text, and HTML output
  • Session history and document restoration
  • Hotkey-based control for hands-free operation
  • Preview, read-aloud, and edit functionality

Use cases

  • Dictating notes or documents directly into a text editor
  • Converting spoken meetings or lectures into structured markdown
  • Generating formatted content from voice instructions for editing

Pros

  • Fully local processing with no cloud dependency
  • Supports real-time streaming transcription and formatting
  • Offers two distinct modes: Global Dictation and Agent Mode
  • Compatible with multiple local LLM servers
  • Free and open-source software

Cons

  • Requires a local LLM server for full functionality on macOS
  • macOS app needs manual installation of dependencies
  • iOS app limited to devices with Apple Intelligence

Frequently asked questions about Speech-to-Markdown

What is Speech-to-Markdown?

Speech-to-Markdown is a macOS menu-bar application and an iPhone/iPad app that converts voice input into structured markdown documents entirely offline. The macOS app uses whisper.cpp for on-device speech recognition and a local LLM server to format the transcribed text into clean markdown, plain text, or HTML in real time.

Who is Speech-to-Markdown for?

The tool is designed for users who prefer hands-free note-taking, writers, developers, and anyone who wants to convert spoken words into structured documents without relying on cloud services or internet connectivity.

Does Speech-to-Markdown require an internet connection?

No, the tool operates entirely offline. The macOS app uses local models and servers, while the iOS app relies on Apple Intelligence, ensuring no data leaves the user's devices.

How do I get started with the macOS app?

Installation is done via a one-line command in Terminal, which sets up dependencies and builds the app. On first launch, grant Microphone and Accessibility permissions, then download a Whisper model from the app's settings.

What are the two modes available in the macOS app?

The macOS app offers Global Dictation, which types transcribed speech directly into any application via a hotkey, and Agent Mode, which streams voice input into a structured document that updates continuously as the user speaks.

What local LLM servers are supported by the macOS app?

The macOS app supports multiple open-source LLM servers, including omlx, Ollama, LM Studio, and llama.cpp, all of which must be running locally on the user's device.

Speech-to-Markdown compared

Reviews