Meta AI enhances machine learning, NLP, and computer vision capabilities.
DocDot
About DocDot
DocDot provides a unified local interface for running multiple PDF parsing tools on macOS. Users can install and switch between parsers such as NanoDoc, PaddleOCR, GLM-OCR, MinerU, and LiteParse without uploading documents to the cloud. The tool supports command-line and web-based interfaces for parsing PDFs, returning output in formats like markdown or JSON. It is designed to integrate directly with agent workflows, including platforms like Codex, Claude Code, OpenClaw, Hermes, and WorkBuddy. DocDot maintains updated versions of supported parsers and allows users to filter tools based on specific capabilities such as multilingual support, table extraction, or chart recognition. The installation process involves a one-line command followed by parser-specific setup, with a web UI available for interactive use.
Key features
- Unified local PDF parsing interface
- Multiple parser options (NanoDoc, PaddleOCR, GLM-OCR, MinerU, LiteParse)
- Command-line and web UI for parsing
- Output in markdown or JSON formats
- Multilingual document support
- Table, formula, and chart extraction
- Agent integration via skills installation
- Automatic parser updates
Use cases
- Extract structured data from complex PDFs for agent workflows
- Process multilingual documents locally without cloud dependency
- Compare parser performance for document processing tasks
Pros
- Local processing without cloud uploads
- Supports multiple OCR and extraction tools
- Command-line and web-based interfaces
- Direct integration with agent platforms
- Updated parser versions maintained automatically
Cons
- Currently limited to M-Chip Macs
- Linux and Windows support not yet available
- Parser selection requires terminal commands
Frequently asked questions about DocDot
What is DocDot and what does it do?
DocDot is a unified local interface for running multiple PDF parsing tools on macOS. It allows users to install and switch between parsers such as NanoDoc, PaddleOCR, GLM-OCR, MinerU, and LiteParse without uploading documents to the cloud.
Who is DocDot designed for?
DocDot is designed for users who need to parse PDFs locally, particularly those working with agent workflows or requiring privacy-focused document processing. It suits developers, researchers, and professionals handling sensitive or large-scale PDF data.
How do I install and set up DocDot?
Installation involves running a one-line command in the terminal, followed by parser-specific setup. Users may need to run 'source ~/.docdot/env' to apply PATH configuration. The process is guided through terminal prompts.
What output formats does DocDot support?
DocDot supports output in markdown and JSON formats, allowing users to integrate parsed data directly into their workflows or agent systems.
Can I use DocDot with agent platforms like Codex or Claude Code?
Yes, DocDot integrates with agent platforms such as Codex, Claude Code, OpenClaw, Hermes, and WorkBuddy. Users can install the DocDot Agent Skill to enable direct interaction with these platforms.
Does DocDot support filtering parsers based on specific capabilities?
Yes, DocDot allows users to filter tools based on capabilities such as multilingual support, table extraction, chart recognition, and more. This helps users select the most suitable parser for their specific document needs.