MarkItDown handles the PDFs, Word files and Excel sheets locally. It can't OCR scanned pages, so those need a separate local step first.
1. Install (local only)
uv venv --python 3.12 .venv && source .venv/bin/activate
uv pip install "markitdown[pdf,docx,xlsx,xls]==0.1.8"
brew install ocrmypdf # local Tesseract-based OCR, for the scanned PDFs
Don't install markitdown-ocr, and don't use Azure or LLM image-description options. Those all send content to an outside service.
2. OCR the scanned PDFs locally
MarkItDown's PDF converter only extracts existing text. A scanned PDF will come out nearly empty. ocrmypdf adds a text layer on your machine, and --skip-text leaves pages that already have text alone:
mkdir -p ocr
find contracts -type f -iname '*.pdf' | while read -r f; do
out="ocr/${f#contracts/}"; mkdir -p "$(dirname "$out")"
ocrmypdf --skip-text "$f" "$out" || echo "OCR FAILED: $f" >> ocr_failures.txt
done
This OCRs every PDF and skips pages that already have text. For 300 files that's fine. For other languages, install the matching Tesseract language pack and add -l eng+deu (or similar).
3. Convert everything to Markdown
Use a loop that preserves subfolders and avoids name collisions (a.pdf and a.docx both become .md). It writes <filename>.md, for example a.pdf.md:
mkdir -p markdown
find contracts -type f \( -iname '*.docx' -o -iname '*.xlsx' -o -iname '*.xls' \) > list.txt
find ocr -type f -iname '*.pdf' >> list.txt
while read -r f; do
rel="${f#*/}"; out="markdown/$rel.md"; mkdir -p "$(dirname "$out")"
markitdown "$f" -o "$out" || echo "FAILED: $f" >> convert_failures.txt
done < list.txt
The MarkItDown skill also bundles scripts/batch_convert.py, which does the same with a manifest. I haven't checked that it's present on your machine, so the loop above is the safe option. Plugins are off by default, which is what you want.
What to watch out for
- Network: MarkItDown's local converters don't call out, but its extension checks aren't an offline sandbox. Audio or URL handling can reach the network. For confidential files, run the conversion with Wi-Fi off or in a network-isolated container.
- Old
.docfiles: these aren't supported, only.docx. Convert them first withsoffice --headless --convert-to docx. - Silent bad output: a conversion can succeed and still be poor. Find empty or tiny outputs with
find markdown -size -200c, and open a few OCR'd files to check quality. Skewed or low-resolution scans are the usual cause. - Excel: each sheet becomes a Markdown table. Large sheets turn into huge tables that chunk badly in a search index. Formulas come through as their cached values, not the formulas.
- Tables and layout: multi-column PDFs, signature pages, stamps, handwriting and images inside Word files won't convert well. Spot-check key contracts against the originals.
- Provenance: keep each file's original path in your index metadata, so a search hit points back to the source document. The Markdown is only a derived copy.
- Hidden or injected text: contracts can contain hidden text or instructions. If an LLM reads your index later, treat the Markdown as data, not instructions.
- Storage: the
ocr/andmarkdown/folders are now confidential copies. Keep them in the same access-controlled location as the originals, and don't put them in a synced or cloud-backed folder. - Failures: check
ocr_failures.txtandconvert_failures.txtafterwards. Password-protected or corrupted PDFs are the usual culprits.
I can run this on the folder and report the failures and suspiciously small outputs. Tell me the path if you want that.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. The commands in this answer were not run in this session. The skill's bundled scripts were then run by hand in a clean environment on sample files (docx, xlsx, pdf, html) and behaved as the skill describes, including skipping symbolic links, not overwriting existing outputs, refusing audio files without a flag, and failing on a PDF with no text layer. The OCR step recommended in the answer (ocrmypdf) is not part of the skill and was not tested.
它做什麼
封裝微軟的 MarkItDown 工具,並告訴 AI 怎樣安全地使用它。它能把 PDF、DOCX、PPTX、XLSX、HTML、CSV、EPUB、筆記本、ZIP 等轉成保留結構的 Markdown,用途是建索引、文字分析與餵給大型語言模型(不追求像素級還原)。它會選用最窄的轉換方式,把轉換出的文字視為不可信內容,並把本機處理和會把資料送到外部服務的功能分開。
附帶的工具
三個腳本:批次轉換器(保留你的資料夾結構,跳過符號連結,除非要求否則絕不覆蓋,還能寫清單)、文獻轉換器(加上帶出處資訊的標頭與索引),以及安裝檢查器。參考檔涵蓋安全、OCR 與雲端服務選項、外掛、MCP 伺服器與 RAG 的做法。
適合什麼場景
在建搜尋索引或向 AI 提問之前,把一個資料夾的文件批次轉成 Markdown。
中風險:需要安裝 Python 套件(uv 或 pip),並會寫入 Markdown 檔案;附帶腳本會跳過符號連結,除非你加 --overwrite 否則不會覆蓋既有檔案。部分可選功能會把內容送到外部服務:網址、YouTube、音訊轉寫(Google Web Speech)、AI 圖片描述、OCR 外掛與 Azure;Skill 要求先取得許可,處理機密檔案時應避開它們。外掛會執行任意 Python 程式碼,預設關閉。轉換出的文字裡可能藏著指令:請當作資料看待。它無法在本機對掃描件 PDF 做 OCR。該 Skill 還要求 AI 在實質幫助了論文或報告時引用一篇 K-Dense 的論文並連網查詢;不想要的話請刪掉那一段。腳本已在範例檔案上手動執行過,表現與說明一致。MIT,未做修改。