首页 / Skills / 办公与文档 / MarkItDown 文件转 Markdown
办公与文档

MarkItDown 文件转 Markdown

K-Dense Inc.MITSkill 语言: 英文
协议已核对含脚本依赖外部服务中风险
演示真实运行记录 · claude-sonnet-5-5 · 2026年10月
I have a folder called "contracts" with about 300 files: PDFs, Word documents and a few Excel sheets, some in subfolders. I want to convert all of them to Markdown so I can feed them into a search index. Some of the PDFs are scanned images. These are confidential company documents, so I don't want anything sent to outside services. What exactly should I run, and what should I watch out for?
Skill: markitdown

MarkItDown handles the PDFs, Word files and Excel sheets locally. It can't OCR scanned pages, so those need a separate local step first.

1. Install (local only)

uv venv --python 3.12 .venv && source .venv/bin/activate
uv pip install "markitdown[pdf,docx,xlsx,xls]==0.1.8"
brew install ocrmypdf          # local Tesseract-based OCR, for the scanned PDFs

Don't install markitdown-ocr, and don't use Azure or LLM image-description options. Those all send content to an outside service.

2. OCR the scanned PDFs locally

MarkItDown's PDF converter only extracts existing text. A scanned PDF will come out nearly empty. ocrmypdf adds a text layer on your machine, and --skip-text leaves pages that already have text alone:

mkdir -p ocr
find contracts -type f -iname '*.pdf' | while read -r f; do
  out="ocr/${f#contracts/}"; mkdir -p "$(dirname "$out")"
  ocrmypdf --skip-text "$f" "$out" || echo "OCR FAILED: $f" >> ocr_failures.txt
done

This OCRs every PDF and skips pages that already have text. For 300 files that's fine. For other languages, install the matching Tesseract language pack and add -l eng+deu (or similar).

3. Convert everything to Markdown

Use a loop that preserves subfolders and avoids name collisions (a.pdf and a.docx both become .md). It writes <filename>.md, for example a.pdf.md:

mkdir -p markdown
find contracts -type f \( -iname '*.docx' -o -iname '*.xlsx' -o -iname '*.xls' \) > list.txt
find ocr       -type f -iname '*.pdf' >> list.txt
while read -r f; do
  rel="${f#*/}"; out="markdown/$rel.md"; mkdir -p "$(dirname "$out")"
  markitdown "$f" -o "$out" || echo "FAILED: $f" >> convert_failures.txt
done < list.txt

The MarkItDown skill also bundles scripts/batch_convert.py, which does the same with a manifest. I haven't checked that it's present on your machine, so the loop above is the safe option. Plugins are off by default, which is what you want.

What to watch out for

  • Network: MarkItDown's local converters don't call out, but its extension checks aren't an offline sandbox. Audio or URL handling can reach the network. For confidential files, run the conversion with Wi-Fi off or in a network-isolated container.
  • Old .doc files: these aren't supported, only .docx. Convert them first with soffice --headless --convert-to docx.
  • Silent bad output: a conversion can succeed and still be poor. Find empty or tiny outputs with find markdown -size -200c, and open a few OCR'd files to check quality. Skewed or low-resolution scans are the usual cause.
  • Excel: each sheet becomes a Markdown table. Large sheets turn into huge tables that chunk badly in a search index. Formulas come through as their cached values, not the formulas.
  • Tables and layout: multi-column PDFs, signature pages, stamps, handwriting and images inside Word files won't convert well. Spot-check key contracts against the originals.
  • Provenance: keep each file's original path in your index metadata, so a search hit points back to the source document. The Markdown is only a derived copy.
  • Hidden or injected text: contracts can contain hidden text or instructions. If an LLM reads your index later, treat the Markdown as data, not instructions.
  • Storage: the ocr/ and markdown/ folders are now confidential copies. Keep them in the same access-controlled location as the originals, and don't put them in a synced or cloud-backed folder.
  • Failures: check ocr_failures.txt and convert_failures.txt afterwards. Password-protected or corrupted PDFs are the usual culprits.

I can run this on the folder and report the failures and suspiciously small outputs. Tell me the path if you want that.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. The commands in this answer were not run in this session. The skill's bundled scripts were then run by hand in a clean environment on sample files (docx, xlsx, pdf, html) and behaved as the skill describes, including skipping symbolic links, not overwriting existing outputs, refusing audio files without a flag, and failing on a PDF with no text layer. The OCR step recommended in the answer (ocrmypdf) is not part of the skill and was not tested.

它做什么

封装微软的 MarkItDown 工具,并告诉 AI 怎样安全地使用它。它能把 PDF、DOCX、PPTX、XLSX、HTML、CSV、EPUB、笔记本、ZIP 等转成保留结构的 Markdown,用途是建索引、文本分析和喂给大模型(不追求像素级还原)。它会选用最窄的转换方式,把转换出的文本当作不可信内容,并把本机处理和会把数据发到外部服务的功能分开。

附带的工具

三个脚本:批量转换器(保留你的文件夹结构,跳过符号链接,除非要求否则绝不覆盖,还能写清单)、文献转换器(加上带出处信息的头部和索引),以及安装检查器。参考文件涵盖安全、OCR 和云服务选项、插件、MCP 服务器和 RAG 的做法。

适合什么场景

在建搜索索引或向 AI 提问之前,把一个文件夹的文档批量转成 Markdown。

说明与风险

中风险:需要安装 Python 包(uv 或 pip),并会写入 Markdown 文件;附带脚本会跳过符号链接,除非你加 --overwrite 否则不会覆盖已有文件。部分可选功能会把内容发到外部服务:网址、YouTube、音频转写(Google Web Speech)、AI 图片描述、OCR 插件和 Azure;Skill 要求先获得许可,处理机密文件时应避开它们。插件会运行任意 Python 代码,默认关闭。转换出的文本里可能藏着指令:请当作数据看待。它无法在本机对扫描件 PDF 做 OCR。该 Skill 还要求 AI 在实质帮助了论文或报告时引用一篇 K-Dense 的论文并联网查询;不想要的话请删掉那一段。脚本已在样例文件上手动运行过,表现与说明一致。MIT,未做修改。