In todays Claude Code adventure. 
When ask why codex catches bugs Claude Code doesn't
(Attachment Link)
That is the correct answer. You can see it while staying in the same model and harness; just ask Claude to do something, then pop up a fresh session and ask it to check; it will usually disagree on some things, possibly finds real bugs. This is valuable, and the computational cost isn't that high - maybe it took 5M tokens to build a semi-large feature; but it only takes maybe 100-200k to read all that code, reason about it, find a bug or two, and give suggestions.
It's funny how the whole language changes in a fresh context. Just like most human programmers, it's somewhat attached to the code it wrote, and tries to explain things away, it literally says "my code". When you commit that to git and ask a fresh agent to modify the same code, it's not "my code" anymore, then it will be quite critical about it, even if you don't ask it to be critical. It's interesting behavior; maybe more than actual emotion of attachment, it's the goal-driven nature, and when it decides its task is done, it requires a bit of pushing to consider the work unfinished again. A fresh context however, hasn't even started, and it's eager to read and comment on code you even hint looking at.
Using different models probably adds a bit more, because their trainings would have different blind spots, different strengths/weaknesses. But probably difference between Opus 4.8 and GPT 5.5 or Claude Code / Codex harness is secondary; just the fresh context itself is the main factor.
Some try to automate this process, but this is also where human-in-the-loop is beneficial. Maybe you can automate the part where a second agent finds a clear bug, and tells the first one to fix it (or fixes it itself), but usually it's not about clear-cut bugs but rather, finding edge cases / design peculiarities you need human to make decisions.
Practical example: a somewhat complex model predictive control electric boiler controller. I let Claude mostly design the implementation details and implement it. Today fixed three bugs:
1) actual logical error in performance optimization: baseline cost from simulating full plan; plans iterated with partial re-simulation (only after the element which changes) - that partial cost compared to original full-plan cost. This is exactly a type of mistake which AI does not seem to do very often - it's quite good at this kind of logical thinking. Happened nevertheless. Good thing - it found it completely by itself; all I had to do is to persuade it that "yes, it really is broken - create artificial tests and look at their results until you fix it". It fixed it.
2) another similar type of logical error - also found it and fixed it without my help beyond persuading it to find it.
3) most interestingly: and this is something Opus 4.8 truly didn't "grok" itself: predicted hot water usage curve is in "liters of hot water per hour". The model correctly subtracts this much of hot liters from the boiler. But
as hot water holds more energy, same liters/hour water usage is modeled as more expensive at boiler temperature of 90degC, compared to 60degC. This drives the algorithm to avoid high temperatures - which reduces spot price arbitrage opportunities ("heat it up to 90degC when electricity is cheap"). Claude says this is correct real-world representation. Many humans would do the same. But what it missed: it missed the mixing valve. If your boiler has 90degC water in it, and you shower with 37degC water,
you are running smaller flow rate of that 90degC water. So liter/hour of "hot water" was the wrong thing to begin with; energy would have been correct. I supplied it with the liters/hour number; it was trained with some typical liters/day values it double-checked against; I can't blame it. 99% of hired human programmers would have done the same mistake; my specification lacked this important detail completely, and it's non-obvious. This is
exactly the case where AI needs a designer who is really into the thing and understands all the subtleties. This is also something that's easily lost when you write the spec in 2 minutes into Jira ticket. But the nice part is: Claude understands what a mixing valve is; it has all the pieces, it just didn't connect them. Very easy to prompt around; just say "don't forget the mixing valve, with 90degC water less flow is needed, energy is all that matters" - and it fixes the code, runs tests, verifies the result.
But you get only this when you actually sit down and do the actual design work - i.e., chat with the implementer (human or AI). Doesn't work if work is managed by passing tickets back and forth. Then the bug remains and causes slight quality regression no one ever notices, because no one was ever interested in that detail, but managing Jira tickets instead and marking them as done.