MarkItDown handles the PDFs, Word files and Excel sheets locally. It can't OCR scanned pages, so those need a separate local step first.
1. Install (local only)
uv venv --python 3.12 .venv && source .venv/bin/activate
uv pip install "markitdown[pdf,docx,xlsx,xls]==0.1.8"
brew install ocrmypdf # local Tesseract-based OCR, for the scanned PDFs
Don't install markitdown-ocr, and don't use Azure or LLM image-description options. Those all send content to an outside service.
2. OCR the scanned PDFs locally
MarkItDown's PDF converter only extracts existing text. A scanned PDF will come out nearly empty. ocrmypdf adds a text layer on your machine, and --skip-text leaves pages that already have text alone:
mkdir -p ocr
find contracts -type f -iname '*.pdf' | while read -r f; do
out="ocr/${f#contracts/}"; mkdir -p "$(dirname "$out")"
ocrmypdf --skip-text "$f" "$out" || echo "OCR FAILED: $f" >> ocr_failures.txt
done
This OCRs every PDF and skips pages that already have text. For 300 files that's fine. For other languages, install the matching Tesseract language pack and add -l eng+deu (or similar).
3. Convert everything to Markdown
Use a loop that preserves subfolders and avoids name collisions (a.pdf and a.docx both become .md). It writes <filename>.md, for example a.pdf.md:
mkdir -p markdown
find contracts -type f \( -iname '*.docx' -o -iname '*.xlsx' -o -iname '*.xls' \) > list.txt
find ocr -type f -iname '*.pdf' >> list.txt
while read -r f; do
rel="${f#*/}"; out="markdown/$rel.md"; mkdir -p "$(dirname "$out")"
markitdown "$f" -o "$out" || echo "FAILED: $f" >> convert_failures.txt
done < list.txt
The MarkItDown skill also bundles scripts/batch_convert.py, which does the same with a manifest. I haven't checked that it's present on your machine, so the loop above is the safe option. Plugins are off by default, which is what you want.
What to watch out for
- Network: MarkItDown's local converters don't call out, but its extension checks aren't an offline sandbox. Audio or URL handling can reach the network. For confidential files, run the conversion with Wi-Fi off or in a network-isolated container.
- Old
.docfiles: these aren't supported, only.docx. Convert them first withsoffice --headless --convert-to docx. - Silent bad output: a conversion can succeed and still be poor. Find empty or tiny outputs with
find markdown -size -200c, and open a few OCR'd files to check quality. Skewed or low-resolution scans are the usual cause. - Excel: each sheet becomes a Markdown table. Large sheets turn into huge tables that chunk badly in a search index. Formulas come through as their cached values, not the formulas.
- Tables and layout: multi-column PDFs, signature pages, stamps, handwriting and images inside Word files won't convert well. Spot-check key contracts against the originals.
- Provenance: keep each file's original path in your index metadata, so a search hit points back to the source document. The Markdown is only a derived copy.
- Hidden or injected text: contracts can contain hidden text or instructions. If an LLM reads your index later, treat the Markdown as data, not instructions.
- Storage: the
ocr/andmarkdown/folders are now confidential copies. Keep them in the same access-controlled location as the originals, and don't put them in a synced or cloud-backed folder. - Failures: check
ocr_failures.txtandconvert_failures.txtafterwards. Password-protected or corrupted PDFs are the usual culprits.
I can run this on the folder and report the failures and suspiciously small outputs. Tell me the path if you want that.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. The commands in this answer were not run in this session. The skill's bundled scripts were then run by hand in a clean environment on sample files (docx, xlsx, pdf, html) and behaved as the skill describes, including skipping symbolic links, not overwriting existing outputs, refusing audio files without a flag, and failing on a PDF with no text layer. The OCR step recommended in the answer (ocrmypdf) is not part of the skill and was not tested.
하는 일
Microsoft의 MarkItDown 도구를 감싸고 AI에게 안전한 사용법을 알려 줍니다. PDF, DOCX, PPTX, XLSX, HTML, CSV, EPUB, 노트북, ZIP 등을 구조를 유지한 Markdown으로 변환하며, 색인 작성, 텍스트 분석, LLM 입력이 용도이고 모양을 똑같이 재현하는 것이 목적은 아닙니다. 가장 범위가 좁은 변환 방식을 고르고, 변환된 텍스트를 신뢰할 수 없는 것으로 다루며, 로컬 작업과 외부 서비스로 데이터를 보내는 기능을 구분합니다.
포함된 도구
스크립트 세 개: 폴더 구조를 유지하고 심볼릭 링크를 건너뛰며 요청이 없으면 덮어쓰지 않고 목록도 쓸 수 있는 일괄 변환기, 출처 정보 머리말과 색인을 붙이는 문헌 변환기, 설치 상태 점검기입니다. 참고 파일은 보안, OCR과 클라우드 옵션, 플러그인, MCP 서버, RAG 방법을 다룹니다.
이런 때 좋습니다
검색 색인을 만들거나 AI에게 물어보기 전에 폴더 안의 문서를 한꺼번에 Markdown으로 바꿀 때.
중간 위험:Python 패키지(uv 또는 pip)를 설치하고 Markdown 파일을 씁니다. 함께 든 스크립트는 심볼릭 링크를 건너뛰고 --overwrite를 주지 않으면 덮어쓰지 않습니다. 일부 선택 기능은 내용을 외부 서비스로 보냅니다(URL, YouTube, 음성 받아쓰기(Google Web Speech), AI 이미지 설명, OCR 플러그인, Azure). 스킬은 먼저 허락을 받고 기밀 파일에는 쓰지 말라고 안내합니다. 플러그인은 임의의 Python을 실행하며 기본은 꺼져 있습니다. 변환된 텍스트에 지시문이 숨어 있을 수 있으니 데이터로만 취급하세요. 스캔된 PDF는 로컬에서 OCR할 수 없습니다. 이 스킬은 논문이나 보고서에 실질적으로 도움이 되었을 때 K-Dense 논문을 인용하고 온라인에서 조회하라고 AI에게 요구합니다. 원치 않으면 해당 부분을 지우세요. 스크립트는 샘플 파일로 직접 실행해 설명대로 동작함을 확인했습니다. MIT, 수정 없음.