홈 / Skills / 개발 생산성 / Grilling 끝까지 질문하기
개발 생산성

Grilling 끝까지 질문하기

Matt PocockMITSkill 언어: 영어
라이선스 확인됨스크립트 없음
데모실행 기록 · claude-opus-5-5 · 2026년 10월
Grill me on this plan before I start building it. We want to move our nightly cron jobs (invoice generation, the email digest, and the CSV data exports) off the single cron server and onto a Redis-backed BullMQ queue, with 3 worker containers on our existing Kubernetes cluster. I'd like to ship it within two weeks.
Skill: grilling

I don't see a codebase in this directory, so this round is based only on your description. If the repo lives somewhere else, point me to it and I'll look up the facts myself.

Round 1. These questions don't depend on each other, so you can answer them in any order.


❓ Q1 - What's actually broken today?: Which of these is driving the move?

  • (a) The cron server is a single point of failure, so if it dies, nothing runs.
  • (b) Jobs take too long or overrun their window and need parallelism.
  • (c) The box is a pet server you want to get rid of.
  • (d) You want retries and visibility into failures.

This decides most of what follows. If it's (a) or (c), a Kubernetes CronJob running the existing scripts unchanged could get you there in 2–3 days, with no Redis and no queue. BullMQ only pays off if you need (b) or (d).

➡️ I suspect it's mostly (a) plus (d). If that's right, a queue is defensible, but you should be able to say why plain k8s CronJobs aren't enough.


❓ Q2 - What triggers the nightly runs?: A queue doesn't schedule anything by itself. Something has to enqueue the jobs at 2am. Options:

  • (a) BullMQ repeatable jobs / job schedulers, defined in code
  • (b) A k8s CronJob that enqueues the jobs and exits
  • (c) The workers enqueue on startup

With (a), the schedule lives in Redis. If you redeploy with a changed cron pattern, the old schedule can linger and fire twice unless you upsert by a stable key.

➡️ (a), using upsertJobScheduler with fixed scheduler IDs, so the schedule is versioned in code and stays idempotent across deploys of all 3 workers.


❓ Q3 - Is invoice generation safe to run twice?: BullMQ guarantees at-least-once delivery, not exactly-once. A worker that gets OOM-killed or evicted mid-job, or misses its lock renewal during a long CPU-bound loop, will have its job marked stalled and picked up by another worker. If invoices get created and emailed twice, customers get billed twice. The same applies to the digest: a duplicate digest is embarrassing but not costly.

What happens today if the invoice script runs twice for the same night?

➡️ Before migrating, make each unit idempotent with a DB uniqueness key such as (customer_id, billing_period) and an "already sent" marker for emails. This is the main thing that has to be in place before launch.


❓ Q4 - Which Redis, and how is it configured?: Is there an existing Redis you plan to share (cache or sessions), or a new one? BullMQ needs maxmemory-policy noeviction. A cache Redis with allkeys-lru will quietly evict job data. Also: is persistence (AOF) turned on, and what happens to the queue if Redis restarts at 2:05am?

➡️ Use a dedicated Redis instance (managed, or a small StatefulSet with AOF), configured with noeviction, and keep it separate from any cache.


❓ Q5 - Where do the CSV exports go?: Containers have ephemeral filesystems. If today's export writes to /var/exports on the cron box, or something reads files from there over SFTP or rsync, that path disappears. Who consumes these files, and how do they get them?

➡️ Write to object storage (S3/GCS) and give consumers a bucket path or signed URL. If an external party pulls the files from the cron box, that's a separate migration to coordinate with them.


❓ Q6 - What does the cron server do that nobody listed?: Old cron boxes usually pick up extra crontab entries, things like log rotation, someone's backup script, or a cleanup find -delete. They also tend to have implicit setup like the timezone (TZ), local credentials, and mounted drives. Have you diffed crontab -l for every user, plus /etc/cron.d, against these three jobs?

➡️ Do a full audit before scoping. Anything left over gets either an owner or an explicit decision to drop it, so it isn't lost when the box is decommissioned.


❓ Q7 - How do you cut over without running jobs twice or zero times?: Options:

  • (a) A hard switch on a given night
  • (b) Shadow mode: the new system runs with side effects disabled and you compare outputs
  • (c) Migrate one job at a time, CSV first, then digest, then invoices last

➡️ (c), starting with the CSV exports because they're lowest-risk, and leaving invoices for last, after a shadow comparison. That also makes two weeks realistic. Moving all three at once, invoices included, in two weeks is the part of the plan I'd push back on hardest.


Answer whichever of these you have opinions on. Your answers, especially to Q1 and Q3, will set up the next round, which covers job granularity, concurrency and rate limits, failure alerting, and worker resource limits.

Real run recorded with claude-code / claude-opus-5-5. Output is shown verbatim, unmodified.

하는 일

계획, 결정, 아이디어에 대해 에이전트가 끈질기게 묻는 면접관이 됩니다. 계획을 결정 트리로 그려 하나의 결정에서 그에 달린 다음 결정으로 뻗어 가며, 암묵적으로 가정한 것이 하나도 남지 않을 때까지 묻습니다.

작동 방식

  1. 지금 바로 답할 수 있는 질문을 번호를 붙여 한 라운드에 모아 내고, 질문마다 추천 답을 함께 제시합니다.
  2. 답을 기다린 뒤, 새로 풀린 질문을 다시 계산해 다음 라운드를 냅니다.
  3. 사실(파일, 도구로 확인할 수 있는 것)은 서브 에이전트가 직접 찾고 사용자에게 묻지 않습니다. 사용자가 정할 것은 '결정'뿐입니다.
  4. 남은 질문이 없을 때 끝나며, 서로 같은 이해에 도달했다고 사용자가 확인하기 전에는 계획을 실행하지 않습니다.

이럴 때 좋습니다

만들기 시작하기 전에 마이그레이션, 아키텍처, 기능 계획을 검증하고 싶을 때.

알아 둘 점

여기 있는 것은 재사용 가능한 질문의 핵심입니다. 작성자의 한 줄짜리 명령 grill-me와 grill-with-docs가 이것을 호출합니다.

참고 및 위험

지시문만 담긴 파일로, 스크립트가 없고 네트워크에 연결하지 않습니다. 사실 확인을 위해 서브 에이전트가 프로젝트 파일을 읽을 수 있습니다. 패키지에는 원본 저장소의 MIT 라이선스(LICENSE)와 agents/openai.yaml(Codex용 표시 이름)도 들어 있습니다.