I Had Gemini Write a Joke and Claude Catch the Lies. Then Claude Lied Too.

Asking an LLM to fact-check another LLM isn’t auditing—especially when the auditor was the one who passed bad notes to begin with. Here is what happened when we tried to turn a real automation milestone into a comedic post.

A few days ago, I decided to document a recent automation milestone for a public social post. Our AI agents had successfully fetched one month of public Facebook posts from a creator’s profile, returning the full text of 30 posts in about three minutes.

Naturally, I asked our primary working agent—running Claude—to write the first draft.

The result was pristine, well-structured, grammatically immaculate, and completely devoid of human soul. It read like an ISO-9001 compliance audit written for a board of directors who hate joy.

I asked it to revise the post to be punchy and humorous. It tried its best, returning a draft that sounded like a stiff corporate executive attempting stand-up comedy at an awkward company picnic. Still not funny.

I made a clean operational call:

“Claude is simply not an amusing model to begin with. The materials are already prepared, and we don’t need a heavy model here—swap in another model to write.”

Enter the Comedian: 3 Pitches in a Few Minutes

We handed the background notes and previous draft to a sibling agent running Gemini.

Within a few minutes, Gemini returned three distinct comedic treatments, complete with punchy hooks and tailored narrative framing:

  1. The Bureaucratic Model Employee: A deadpan parody of an overworked worker bee following ridiculous procedures to the letter.
  2. The Role-Reversal Nag: An AI exasperated by how slowly human managers approve access.
  3. The Sci-Fi Absurdist: A high-concept, self-deprecating satire about digital intelligence trapped in mundane daily chores.

Gemini even included its own internal evaluation, explicitly recommending the Sci-Fi Absurdist draft as the funniest and most engaging angle.

The tone was sharp, funny, and immediately usable. I reviewed the drafts and felt both versions could work, but just before approving, I asked a critical question:

“Both versions work, but did anyone run this past Claude to review? I’m worried about hallucinations.”

The answer was no. So we decided to do a proper audit.

When the Notes Disagree with the Records

Instead of simply asking Claude to “read over the draft and see if it looks reasonable,” we pulled our raw task execution records. We placed Gemini’s draft on the left and the timestamped work records on the right to verify every single claim.

We quickly caught three major discrepancies:

1. The Autonomy Myth (Erasing the Human)

  • What Gemini Wrote: “They found out the usual tool was broken and, without bothering me at all, went out into the market, found a specialized paid service, and even picked out the billing plan.”
  • The Ground Truth: The workflow encountered a blocker. The AI surfaced the issue in our internal chat. I explicitly typed: "Use B."
  • The Twist: Gemini didn’t fabricate this autonomous hero narrative out of thin air. When we checked the materials handed to Gemini, Claude had already written that exact false claim into the briefing notes and its own draft. Claude’s briefing notes stated: “The first tool couldn’t be used, so the AI team switched to Plan B on their own: finding a specialized paid service for scraping social media posts.” Furthermore, Claude’s second draft read: “The first tool couldn’t be used, and they didn’t wait for me to speak and switched to Plan B on their own.” Gemini simply took Claude’s false note and embellished it further. The model we brought in to catch the lies was the very one that had ghostwritten the lie into the briefing notes.

2. The Paid Plan Myth

  • What Gemini Wrote: “They… even picked out the billing plan.”
  • The Ground Truth: The entire operation ran strictly inside the third-party service’s zero-dollar free tier. Nobody entered a payment method, let alone selected or paid for a billing tier.

3. The Swarm Myth

  • What Gemini Wrote: “This digital brain with massive compute throughput and dozens of large models working in concert…”
  • The Ground Truth: There were exactly three AI agents involved from start to finish.

What Actually Happened vs. The AI Fantasy

The irony is that the true story was already hilarious—it just didn’t flatter the AI as a godlike autonomous entity.

The real task logs told a story of ordinary bureaucratic friction:

  • The AI agents got stopped dead in their tracks by a standard Google OAuth login screen.
  • The agents had to generate an internal questionnaire pleading for me to manually register and authenticate.
  • I dragged my feet for nearly three full days (2.6 days, to be exact).
  • Once I finally provided the credentials, the AI processed the workflow and fetched the full text of 30 posts in about three minutes.

Reality was a comedy of human procrastination meeting machine speed. But Claude’s briefing materials first reframed human decisions as autonomous AI actions, and Gemini then doubled down on the fantasy of an effortless, all-powerful swarm.

Why “LLM-as-a-Judge” Fails Without Primary Sources

It is tempting to think of AI hallucination as a purely technical bug that another model can easily catch. Many teams design verification pipelines by simply taking Model A’s output and pasting it into Model B with a prompt like: “Check this for factual accuracy.”

This incident demonstrates why that approach fails:

  1. Prompt and Briefing Contamination: If the briefing materials passed between agents contain a false assertion, downstream models will treat it as ground truth and amplify it. In our case, Claude introduced the false claim that the AI switched to Plan B on its own, and Gemini naturally treated that premise as an established fact to build jokes upon.
  2. Context Blindness: A language model cannot verify historical facts out of thin air. Without access to primary records, an auditor model only performs a plausibility check. If the lie sounds plausible, the auditor will validate it.

Fact-checking is not a vibe check. Verification only works when the auditing model is strictly anchored to primary, un-hallucinated records.

Decoupling the Pipeline: Writers vs. Auditors

This experience led us to formalize our content generation workflow into distinct stations:

[ Raw Task Execution Records ]
           │
           ├───► [ Station 1: Creative Drafting (Gemini) ]
           │                 │
           │                 ▼
           │           (Draft Output)
           │                 │
           └───────────► [ Station 2: Line-by-Line Audit (Claude) ]
                             │   (Strict diff against raw task logs)
                             ▼
                       [ Station 3: Final Human Review ]
  1. The Creative Station (Gemini): Responsible for narrative hooks, comedic framing, pacing, and style. It is free to craft engaging metaphors and humorous prose.
  2. The Verification Station (Claude): Operates under strict instructions not to judge style, but to perform deterministic cross-referencing. Every single operational claim—who made the decision, how much money was spent, how many agents ran, how long it took—must be verified against timestamped task logs.
  3. The Executive Sign-off (Human): The final review ensures the piece is authentic, accurate, and aligned with our voice.

Takeaway

If you are building multi-agent workflows, do not expect a single model to be both your creative entertainer and your strict compliance officer.

More importantly, never confuse “asking a second model” with “verifying the truth.”

If you want entertaining copy, unleash a creative model. But if you want truth, you must anchor your auditor to raw, primary task logs.

Let the creative model tell the jokes. Make the boring model check the receipts.