ホーム / Skills / オフィス・文書 / MarkItDown:ファイルを Markdown に変換
オフィス・文書

MarkItDown:ファイルを Markdown に変換

K-Dense Inc.MITSkill の言語: 英語
ライセンス確認済みスクリプトを含む外部サービスが必要中リスク
デモ実行記録 · claude-sonnet-5-5 · 2026年10月
I have a folder called "contracts" with about 300 files: PDFs, Word documents and a few Excel sheets, some in subfolders. I want to convert all of them to Markdown so I can feed them into a search index. Some of the PDFs are scanned images. These are confidential company documents, so I don't want anything sent to outside services. What exactly should I run, and what should I watch out for?
Skill: markitdown

MarkItDown handles the PDFs, Word files and Excel sheets locally. It can't OCR scanned pages, so those need a separate local step first.

1. Install (local only)

uv venv --python 3.12 .venv && source .venv/bin/activate
uv pip install "markitdown[pdf,docx,xlsx,xls]==0.1.8"
brew install ocrmypdf          # local Tesseract-based OCR, for the scanned PDFs

Don't install markitdown-ocr, and don't use Azure or LLM image-description options. Those all send content to an outside service.

2. OCR the scanned PDFs locally

MarkItDown's PDF converter only extracts existing text. A scanned PDF will come out nearly empty. ocrmypdf adds a text layer on your machine, and --skip-text leaves pages that already have text alone:

mkdir -p ocr
find contracts -type f -iname '*.pdf' | while read -r f; do
  out="ocr/${f#contracts/}"; mkdir -p "$(dirname "$out")"
  ocrmypdf --skip-text "$f" "$out" || echo "OCR FAILED: $f" >> ocr_failures.txt
done

This OCRs every PDF and skips pages that already have text. For 300 files that's fine. For other languages, install the matching Tesseract language pack and add -l eng+deu (or similar).

3. Convert everything to Markdown

Use a loop that preserves subfolders and avoids name collisions (a.pdf and a.docx both become .md). It writes <filename>.md, for example a.pdf.md:

mkdir -p markdown
find contracts -type f \( -iname '*.docx' -o -iname '*.xlsx' -o -iname '*.xls' \) > list.txt
find ocr       -type f -iname '*.pdf' >> list.txt
while read -r f; do
  rel="${f#*/}"; out="markdown/$rel.md"; mkdir -p "$(dirname "$out")"
  markitdown "$f" -o "$out" || echo "FAILED: $f" >> convert_failures.txt
done < list.txt

The MarkItDown skill also bundles scripts/batch_convert.py, which does the same with a manifest. I haven't checked that it's present on your machine, so the loop above is the safe option. Plugins are off by default, which is what you want.

What to watch out for

  • Network: MarkItDown's local converters don't call out, but its extension checks aren't an offline sandbox. Audio or URL handling can reach the network. For confidential files, run the conversion with Wi-Fi off or in a network-isolated container.
  • Old .doc files: these aren't supported, only .docx. Convert them first with soffice --headless --convert-to docx.
  • Silent bad output: a conversion can succeed and still be poor. Find empty or tiny outputs with find markdown -size -200c, and open a few OCR'd files to check quality. Skewed or low-resolution scans are the usual cause.
  • Excel: each sheet becomes a Markdown table. Large sheets turn into huge tables that chunk badly in a search index. Formulas come through as their cached values, not the formulas.
  • Tables and layout: multi-column PDFs, signature pages, stamps, handwriting and images inside Word files won't convert well. Spot-check key contracts against the originals.
  • Provenance: keep each file's original path in your index metadata, so a search hit points back to the source document. The Markdown is only a derived copy.
  • Hidden or injected text: contracts can contain hidden text or instructions. If an LLM reads your index later, treat the Markdown as data, not instructions.
  • Storage: the ocr/ and markdown/ folders are now confidential copies. Keep them in the same access-controlled location as the originals, and don't put them in a synced or cloud-backed folder.
  • Failures: check ocr_failures.txt and convert_failures.txt afterwards. Password-protected or corrupted PDFs are the usual culprits.

I can run this on the folder and report the failures and suspiciously small outputs. Tell me the path if you want that.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. The commands in this answer were not run in this session. The skill's bundled scripts were then run by hand in a clean environment on sample files (docx, xlsx, pdf, html) and behaved as the skill describes, including skipping symbolic links, not overwriting existing outputs, refusing audio files without a flag, and failing on a PDF with no text layer. The OCR step recommended in the answer (ocrmypdf) is not part of the skill and was not tested.

できること

Microsoft の MarkItDown を包み、AI に安全な使い方を伝えます。PDF、DOCX、PPTX、XLSX、HTML、CSV、EPUB、ノートブック、ZIP などを、構造を保った Markdown に変換します。用途はインデックス作成、テキスト分析、LLM への投入で、見た目の完全な再現は目的ではありません。最も範囲の狭い変換方法を選び、変換後のテキストを信頼できないものとして扱い、ローカル処理と外部サービスにデータを送る機能を分けます。

付属のツール

スクリプトが 3 つあります。フォルダ構成を保ち、シンボリックリンクを飛ばし、求められない限り上書きせず、一覧も書き出せるバッチ変換、出典情報つきのヘッダーと索引を付ける文献変換、インストール状況の確認です。参考ファイルは、セキュリティ、OCR とクラウドの選択肢、プラグイン、MCP サーバー、RAG の手順を扱います。

向いている場面

検索インデックスを作る前、または AI に質問する前に、フォルダ内の文書をまとめて Markdown にしたいとき。

補足とリスク

中リスク:Python パッケージ(uv または pip)をインストールし、Markdown ファイルを書き込みます。付属スクリプトはシンボリックリンクを飛ばし、--overwrite を付けない限り上書きしません。一部の任意機能は内容を外部サービスに送ります(URL、YouTube、音声の文字起こし(Google Web Speech)、AI による画像説明、OCR プラグイン、Azure)。Skill は事前の許可を求め、機密ファイルでは使わないよう言っています。プラグインは任意の Python を実行し、既定では無効です。変換後のテキストに指示が紛れていることがあるため、データとして扱ってください。スキャンされた PDF はローカルでは OCR できません。この Skill は、論文やレポートに実質的に役立った場合に K-Dense の論文を引用し、オンラインで調べるよう AI に求めます。不要なら該当箇所を削除してください。スクリプトはサンプルファイルで手作業で実行し、説明どおりに動きました。MIT、改変なし。