Home / Skills / Research & Learning / Doublecheck (AI output verification)
Research & Learning

Doublecheck (AI output verification)

awesome-copilot communityMITSkill language: English
License verifiedNo scriptsNeeds external serviceMedium risk
DemoReal run · claude-sonnet-5-5 · October 2026
Using the doublecheck skill in one-shot mode, verify this fictional AI-written paragraph: "Python was created by Guido van Rossum and first released in 1991. According to the 2019 Henderson v. Alder Systems decision of the 9th Circuit, open-source maintainers are personally liable for downstream security flaws. A 2023 survey found that 87.4% of enterprises have fully migrated to Python 3. The GPL requires derivative works to be distributed under the same license."
Skill: doublecheck
Read: assets/verification-report-template.md

Verification Report

Heads up: I couldn't find support for two claims, and one of them contradicts well-established law. The Henderson v. Alder Systems citation, its holding, and the "87.4%" statistic should not be relied on without a primary source.

Limits of this run: I have no web search tool in this session, so Layer 2 was not run live. Ratings come from my own knowledge plus the adversarial review. The links below are where to check, and I haven't fetched them. Because of that, I rated nothing VERIFIED.

Summary

Text verified: a four-sentence paragraph on Python, a court decision, a migration survey, and the GPL
Claims extracted: 6

Rating Count
VERIFIED 0
PLAUSIBLE 3
UNVERIFIED 0
DISPUTED 1
FABRICATION RISK 2

Items requiring attention: 3


Flagged Items (Review These First)

C3 – "Henderson v. Alder Systems" (9th Cir. 2019)

  • Claim: "According to the 2019 Henderson v. Alder Systems decision of the 9th Circuit…"
  • Rating: FABRICATION RISK
  • Finding: I don't recall any such case. It has the shape of a fabricated citation: a plausible party name, a court, and a year, with no reporter cite or docket number. I haven't searched a legal database to confirm it doesn't exist.
  • Source: Search it yourself at https://www.courtlistener.com/ and https://scholar.google.com/ (case law), and at the 9th Circuit opinions page, https://www.ca9.uscourts.gov/opinions/
  • Recommendation: Treat it as nonexistent until you find it in Westlaw, Lexis, or CourtListener. Don't cite it.

C4 – Maintainers "personally liable for downstream security flaws"

  • Claim: Open-source maintainers are personally liable for downstream security flaws.
  • Rating: DISPUTED (also FABRICATION RISK, since it depends on C3)
  • Finding: This contradicts the general legal picture. Open-source licenses such as MIT, Apache-2.0, and GPL disclaim warranties and liability. I know of no US appellate rule imposing blanket personal liability on maintainers. The EU Cyber Resilience Act debate is about this question, and the final text deliberately limits obligations for non-commercial open-source contributors. The claim is also a sweeping overgeneralization, since liability would depend on jurisdiction, facts, and commercial involvement.
  • Source: https://www.gnu.org/licenses/gpl-3.0.html (sections 15–16, disclaimer of warranty and limitation of liability); https://opensource.org/license/mit; and for the EU, https://digital-strategy.ec.europa.eu/en/policies/cyber-resilience-act
  • Recommendation: Remove it, or replace it with a sourced, jurisdiction-specific statement from a lawyer.

C5 – "87.4% of enterprises have fully migrated to Python 3"

  • Claim: A 2023 survey found that 87.4% of enterprises have fully migrated to Python 3.
  • Rating: FABRICATION RISK
  • Finding: The survey is unnamed, and the figure is precise and unsourced. "Fully migrated" is also a strong and hard-to-measure claim. Python 2 reached end of life on 1 January 2020. Python Developers Survey results (JetBrains and PSF) do track Python 3 usage, but I'm not aware of this specific number.
  • Source: https://www.python.org/doc/sunset-python-2/ (end of life); https://lp.jetbrains.com/python-developers-survey-2023/ (nearest real survey)
  • Recommendation: Drop the number unless you can find the survey. If you find one, check that it says "enterprises" and "fully migrated."

All Claims

VERIFIED

None. See the limits note above.

PLAUSIBLE

C1 – Creator

C2 – First release in 1991

C6 – GPL and derivative works

  • Claim: The GPL requires derivative works to be distributed under the same license.
  • Notes: This is broadly right, but it is missing qualifiers (pattern 7). The copyleft obligation applies when you convey or distribute a covered work, not when you modify it privately. GPLv2 and GPLv3 differ in detail. LGPL and linking exceptions differ again. Whether something counts as a "derivative work" (for example, through dynamic linking) is contested. Source: https://www.gnu.org/licenses/gpl-faq.html

UNVERIFIED

None.

DISPUTED

C4. See the flagged items.

FABRICATION RISK

C3 and C5. See the flagged items.


Internal Consistency

No contradictions within the text. The paragraph does mix accurate background (C1, C2, C6) with unsupported claims, which makes the unsupported ones look more credible.


What Was Not Checked

  • No live web search or legal database lookup was run, so I haven't confirmed that the case in C3 doesn't exist.
  • I haven't read the text of the GPL or the PSF survey data directly.
  • C6's "derivative work" boundary is a legal question that needs a lawyer.

Limitations

  • This tool accelerates human verification; it does not replace it.
  • Web search results may not include the most recent information or paywalled sources, and none were used here.
  • The adversarial review uses the same underlying model that may have produced the original output. It catches many issues but cannot catch all of them.
  • A claim rated VERIFIED means a supporting source was found, not that the claim is definitely correct. Sources can be wrong too.
  • Claims rated PLAUSIBLE may still be wrong. The absence of contradicting evidence is not proof of accuracy.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so nothing was fetched from the web and no file was written. The example is fictional. No web search was enabled, so the model says it could not run the live source checks and rated no claim as verified.

What it does

Runs a three-layer verification on AI output. Layer 1 extracts claims (factual, statistical, citation, entity, causal, temporal) and checks internal consistency. Layer 2 searches the web for supporting or contradicting sources and gives you links. Layer 3 reviews the text adversarially against a checklist of hallucination patterns such as fabricated citations, precise numbers with no source, plausible-but-wrong associations and missing qualifiers. Each claim is rated VERIFIED, PLAUSIBLE, UNVERIFIED, DISPUTED or FABRICATION RISK.

How it works

  1. In one-shot mode you give the text, or point to earlier output, and get a full report.
  2. In active mode it stays on for the conversation and adds a short verification section to substantive answers, escalating to a full report when a claim looks wrong or the content is legal or regulatory.
  3. You can switch it off any time by saying so.

Good for

Reviewing summaries, research notes and drafts that cite cases, statistics or standards before you rely on them.

Notes & risks

Pure instructions: no scripts and no file writes. Layer 2 needs a web search tool in your agent; the claims from your text become search queries, so do not check confidential text. Without web access the skill says it could not run live checks. The skill aims to give you sources, not verdicts: VERIFIED only means a supporting source was found, the adversarial pass uses the same kind of model, and nothing here replaces a lawyer, doctor or other expert. Active mode stays on until you turn it off.