High-Quality Few Shot Code Generation With Claude Code and Codex
Achieve high quality code in a short series of rounds
I setup Codex Cloud a few months ago to perform adversarial reviews of diffs during CI. Codex surfaces bugs, Claude Code addresses the bugs, it gets reviewed by Codex again, then I merge on clean CI-run and a Codex thumbs up reaction on the pull request. This loop, depending on the complexity of the task, results in high-quality code and it works really well, but it often takes ~7-10 rounds.
I decided to implement a skill that contains my preferred mechanics of how Claude Code interacts with Github, Codex reviews, approvals, etc. But I changed one thing in this skill: I implemented a 3 round constraint on the number of review rounds up-front and the LLM is aware of the constraints and the objectives. The result: fewer loops.
The best part: I'm using Opus 5 with Extra effort. Not Fable, not Fable 5.1, not Ultracode, not Sol 5.6 Ultra. Opus 5 Extra is a middle of the road reasoning effort on a daily-driver LLM.
The results I am seeing are really good. Here's the skill in full:
---
name: codex-pr-loop
description: "Drive a GitHub PR through Codex's automated review until sign-off, under a caller-specified round budget and a strict one-push-per-review-round cost invariant: snapshot findings, fix and validate the whole batch locally, push exactly once, reply on and resolve each thread, and iterate until Codex leaves a thumbs-up. Use when a PR has Codex review comments, or the user says \"address the codex comments\", \"get this PR through codex review\", \"codex flagged some things\"."
---
# Codex PR review loop
Work a pull request through rounds of Codex review until Codex signals approval with a 👍 reaction on the PR. Every remote update is expensive — it triggers a Codex review cycle and a CI cycle — so the loop is built around one push per round, inside a bounded number of rounds.
## Hard rules — never violate these
1. **Never write `@codex`** — not in a review reply, a PR comment, a commit message, a PR description, or a branch name. Refer to it as "Codex" in prose only. An at-mention kicks off an unwanted run, and once pushed it can't be taken back.
2. **One push per review round.** See the invariant below. This is the rule the whole skill exists to enforce.
3. **Never exceed the round budget.** See below. Budget exhaustion is a stop-and-report, never one more push.
4. **Never force-push, never amend a published commit.** Fixes are new commits on top; history stays append-only so Codex's diff-since-last-review stays meaningful.
5. **Don't stop early.** Within budget, keep looping until Codex leaves 👍 on the PR or a stop condition fires. A round with no new findings is not sign-off — only the 👍 is.
6. **Fix, don't silence.** Never resolve a thread that wasn't actually addressed in code, and never suppress a lint/check to make a finding go away.
## Round budget — set by the caller
The total number of review rounds is **specified at invocation**, not chosen by the skill:
- If the user stated a budget ("max 3 rounds", "one round only", "up to 5 pushes"), use it.
- If they didn't, **ask before round 1** — a single question, alongside confirming which PR.
- Only when running unattended with nobody to ask (a scheduled or headless run), default to **3 rounds** and say so explicitly in the first round summary.
Track the round number from 1. When the budget is exhausted, stop and report — do not start another round, and do not push. The user can re-invoke with a fresh budget after reading where things stand. If the user says "keep going until it's signed off" with no number, treat that as an explicit unbounded budget and confirm it once, in writing, in the round-1 summary.
The same-finding-three-rounds rule below is a separate circuit breaker, not the budget — both apply.
## Cost-control invariant: one push per review round
A **review round** is the complete set of Codex findings associated with one PR head commit.
1. Snapshot every open Codex finding for the current head, merged with the deferred ledger (below).
2. Fix and validate the entire set locally. Local iteration is unlimited — as many edits and local commits as needed.
3. Immediately before pushing, refresh the review state **exactly once**. Incorporate any additional findings associated with that same head. Findings that appear after this refresh are deferred to the next round; do not refresh a second time.
4. Push exactly once, after the pre-push gate passes.
5. After that push, the round is **locked**. Do not push again until the lock releases.
Never push a quick follow-up for a missed edit, formatting issue, failing check, or newly noticed problem from the same finding set. Record it in the deferred ledger and include it in the next fully batched round.
Never split one review round across multiple remote updates. The purpose of this rule is to trigger at most one Codex review cycle and one CI cycle per set of findings.
### Lock release
The round lock releases when CI for the pushed head has reached a terminal state **and** any one of:
- Codex has reviewed the pushed head (new findings → next round, if budget remains), or
- Codex signed off (👍 → done), or
- the response timeout is reached (default: 30 minutes of polling → stop and report).
**Deadlock escape:** if CI reaches a terminal *failing* state on the pushed head, the lock releases immediately without waiting for Codex — some setups skip review on red builds. Fold the CI failures into the next round's finding set and proceed, budget permitting.
### Push failure
If a push command reports failure, query the remote head before retrying:
```bash
git ls-remote origin refs/heads/<branch>
```
Retry only when the remote ref is provably unchanged from the round's starting head. If the remote moved, treat the round's push as consumed and stop to reassess — report to the user rather than racing.
### Deferred ledger
Anything noticed after the refresh cutoff, or after the push, goes in `.git/codex-deferred.md` — inside `.git/`, so it is structurally uncommittable and survives across shell calls.
```bash
echo "- [round <n>] <file:line> — <what needs doing>" >> .git/codex-deferred.md
```
Step 1 of every round **must** read this file and merge its entries into the round's finding set, then clear the entries it took. A deferred item that isn't merged forward is a silent regression. Anything still in the ledger when the loop stops goes in the final report.
### Pre-push gate
Push only when every line is true. If any fails, fix it locally — do not push and follow up.
- [ ] The round budget has not been exhausted (this round is within it).
- [ ] Every round finding has a disposition (fixed / declined / needs-user).
- [ ] All intended commits exist locally.
- [ ] Staged and committed diffs contain no unrelated work.
- [ ] Required local checks pass.
- [ ] No commit message, and no staged content, contains the string `@codex` — verify: `git log origin/<branch>..HEAD --format=%B | grep -i '@codex'` and `git diff origin/<branch>..HEAD | grep -i '@codex'` both return nothing.
- [ ] The remote head still equals the round's starting head.
- [ ] No push has already occurred for this round.
## Setup (once per PR)
```bash
gh pr view --json number,url,headRefName,baseRefName,state,isDraft
```
Set `OWNER`, `REPO`, `PR` from that, and establish the round budget. Record the round's starting head: `git rev-parse origin/<branch>`.
Identify the Codex bot's login rather than assuming it — read it off the existing reviews:
```bash
gh api graphql -f query='
query($owner:String!,$repo:String!,$pr:Int!){
repository(owner:$owner,name:$repo){ pullRequest(number:$pr){
reviews(last:20){ nodes{ author{login} state submittedAt } } } } }' \
-F owner=$OWNER -F repo=$REPO -F pr=$PR
```
The reviewer login that looks like the Codex integration (e.g. a `...codex...` app/bot account) is `CODEX_LOGIN`. If several candidates appear, ask the user which one before proceeding.
## Round procedure
### 1. Snapshot the findings
```bash
gh api graphql -f query='
query($owner:String!,$repo:String!,$pr:Int!){
repository(owner:$owner,name:$repo){ pullRequest(number:$pr){
reviewThreads(first:100){ nodes{
id isResolved isOutdated path line
comments(first:20){ nodes{ databaseId body author{login} createdAt } } } } } } }' \
-F owner=$OWNER -F repo=$REPO -F pr=$PR
```
Take threads where `isResolved == false` and the first comment's author is `CODEX_LOGIN`. Keep each thread's `id` (for resolving) and its first comment's `databaseId` (for replying). Also read the review body for PR-level findings, and merge in `.git/codex-deferred.md`.
Triage each finding into **fix**, **decline** (wrong, out of scope, or a deliberate choice), or **needs-user** (a product/design decision you shouldn't make alone).
### 2. Fix and validate the whole batch locally
Make every "fix" change before touching the remote. Run the repo's own checks (tests, typecheck, lint — whatever `CONTRIBUTING`/`package.json`/`Makefile` defines) until they pass. Iterate freely here; nothing leaves the machine yet.
### 3. Commit one per finding
```bash
git add <files for this finding>
git commit -m "fix: <what changed, in your own words>"
```
Descriptive of the change itself. Do not paste Codex's comment text and do not mention it by handle.
### 4. Refresh once, run the gate, push once
Re-run the step-1 query. Fold in any new findings on the same head (repeat steps 2–3 for them). Then run the pre-push gate, then:
```bash
git push
```
Record each finding's SHA: `git log --oneline -n <count>`. The round is now locked and one unit of budget is consumed.
### 5. Reply, then resolve
For each **fixed** finding:
```bash
gh api repos/$OWNER/$REPO/pulls/$PR/comments/<databaseId>/replies \
-f body='Fixed in <sha> — <one line on what changed and why it addresses this>.'
gh api graphql -f query='mutation($id:ID!){ resolveReviewThread(input:{threadId:$id}){ thread{ isResolved } } }' \
-F id=<threadId>
```
For **declined** and **needs-user** findings: reply with the reasoning and **leave the thread unresolved** — a human decides those.
### 6. Post a round summary
One PR-level comment per round, so the trail is readable without opening every thread:
```bash
gh pr comment $PR --body "$(cat <<'EOF'
### Review round <n> of <budget>
| Finding | Action | Commit |
|---|---|---|
| <file:line — short description> | Fixed | <sha> |
| <file:line — short description> | Declined — <reason> | — |
Checks: <what you ran and the result>
Deferred to next round: <items, or none>
EOF
)"
```
### 7. Wait for lock release
Poll; don't busy-loop. `sleep 120` between checks, a few checks per shell call.
```bash
# CI state for the pushed head
gh pr view $PR --json statusCheckRollup
# thumbs-up on the PR itself
gh api repos/$OWNER/$REPO/issues/$PR/reactions \
--jq '[.[] | select(.content=="+1") | .user.login]'
# new review activity
gh api graphql -f query='
query($owner:String!,$repo:String!,$pr:Int!){
repository(owner:$owner,name:$repo){ pullRequest(number:$pr){
reviews(last:5){ nodes{ author{login} state submittedAt body } }
reviewThreads(first:100){ nodes{ id isResolved comments(first:1){ nodes{ author{login} createdAt } } } } } } }' \
-F owner=$OWNER -F repo=$REPO -F pr=$PR
```
When `CODEX_LOGIN` appears in the 👍 reactions, stop — the PR has sign-off. Otherwise, on lock release, start the next round at step 1 with a fresh starting head, if budget remains.
## Stop and report to the user when
- Codex has left 👍 → report the rounds it took and anything left declined/unresolved.
- **The round budget is exhausted** → report state; do not start another round.
- Codex raises the **same** finding for a third round → the fix isn't landing; describe the disagreement.
- A finding needs a product or design decision.
- The remote head moved under you and the round's push was consumed.
- The response timeout (~30 min) elapsed with no Codex activity.
In every case: rounds used out of budget, which findings were fixed (with SHAs), which were declined and why, which threads are still open, and what's sitting in the deferred ledger.