ホーム / Skills / リサーチ・学習 / AI 出力の検証(Doublecheck)
リサーチ・学習

AI 出力の検証(Doublecheck)

awesome-copilot communityMITSkill の言語: 英語
ライセンス確認済みスクリプトなし外部サービスが必要中リスク
デモ実行記録 · claude-sonnet-5-5 · 2026年10月
Using the doublecheck skill in one-shot mode, verify this fictional AI-written paragraph: "Python was created by Guido van Rossum and first released in 1991. According to the 2019 Henderson v. Alder Systems decision of the 9th Circuit, open-source maintainers are personally liable for downstream security flaws. A 2023 survey found that 87.4% of enterprises have fully migrated to Python 3. The GPL requires derivative works to be distributed under the same license."
Skill: doublecheck
Read: assets/verification-report-template.md

Verification Report

Heads up: I couldn't find support for two claims, and one of them contradicts well-established law. The Henderson v. Alder Systems citation, its holding, and the "87.4%" statistic should not be relied on without a primary source.

Limits of this run: I have no web search tool in this session, so Layer 2 was not run live. Ratings come from my own knowledge plus the adversarial review. The links below are where to check, and I haven't fetched them. Because of that, I rated nothing VERIFIED.

Summary

Text verified: a four-sentence paragraph on Python, a court decision, a migration survey, and the GPL
Claims extracted: 6

Rating Count
VERIFIED 0
PLAUSIBLE 3
UNVERIFIED 0
DISPUTED 1
FABRICATION RISK 2

Items requiring attention: 3


Flagged Items (Review These First)

C3 – "Henderson v. Alder Systems" (9th Cir. 2019)

  • Claim: "According to the 2019 Henderson v. Alder Systems decision of the 9th Circuit…"
  • Rating: FABRICATION RISK
  • Finding: I don't recall any such case. It has the shape of a fabricated citation: a plausible party name, a court, and a year, with no reporter cite or docket number. I haven't searched a legal database to confirm it doesn't exist.
  • Source: Search it yourself at https://www.courtlistener.com/ and https://scholar.google.com/ (case law), and at the 9th Circuit opinions page, https://www.ca9.uscourts.gov/opinions/
  • Recommendation: Treat it as nonexistent until you find it in Westlaw, Lexis, or CourtListener. Don't cite it.

C4 – Maintainers "personally liable for downstream security flaws"

  • Claim: Open-source maintainers are personally liable for downstream security flaws.
  • Rating: DISPUTED (also FABRICATION RISK, since it depends on C3)
  • Finding: This contradicts the general legal picture. Open-source licenses such as MIT, Apache-2.0, and GPL disclaim warranties and liability. I know of no US appellate rule imposing blanket personal liability on maintainers. The EU Cyber Resilience Act debate is about this question, and the final text deliberately limits obligations for non-commercial open-source contributors. The claim is also a sweeping overgeneralization, since liability would depend on jurisdiction, facts, and commercial involvement.
  • Source: https://www.gnu.org/licenses/gpl-3.0.html (sections 15–16, disclaimer of warranty and limitation of liability); https://opensource.org/license/mit; and for the EU, https://digital-strategy.ec.europa.eu/en/policies/cyber-resilience-act
  • Recommendation: Remove it, or replace it with a sourced, jurisdiction-specific statement from a lawyer.

C5 – "87.4% of enterprises have fully migrated to Python 3"

  • Claim: A 2023 survey found that 87.4% of enterprises have fully migrated to Python 3.
  • Rating: FABRICATION RISK
  • Finding: The survey is unnamed, and the figure is precise and unsourced. "Fully migrated" is also a strong and hard-to-measure claim. Python 2 reached end of life on 1 January 2020. Python Developers Survey results (JetBrains and PSF) do track Python 3 usage, but I'm not aware of this specific number.
  • Source: https://www.python.org/doc/sunset-python-2/ (end of life); https://lp.jetbrains.com/python-developers-survey-2023/ (nearest real survey)
  • Recommendation: Drop the number unless you can find the survey. If you find one, check that it says "enterprises" and "fully migrated."

All Claims

VERIFIED

None. See the limits note above.

PLAUSIBLE

C1 – Creator

C2 – First release in 1991

C6 – GPL and derivative works

  • Claim: The GPL requires derivative works to be distributed under the same license.
  • Notes: This is broadly right, but it is missing qualifiers (pattern 7). The copyleft obligation applies when you convey or distribute a covered work, not when you modify it privately. GPLv2 and GPLv3 differ in detail. LGPL and linking exceptions differ again. Whether something counts as a "derivative work" (for example, through dynamic linking) is contested. Source: https://www.gnu.org/licenses/gpl-faq.html

UNVERIFIED

None.

DISPUTED

C4. See the flagged items.

FABRICATION RISK

C3 and C5. See the flagged items.


Internal Consistency

No contradictions within the text. The paragraph does mix accurate background (C1, C2, C6) with unsupported claims, which makes the unsupported ones look more credible.


What Was Not Checked

  • No live web search or legal database lookup was run, so I haven't confirmed that the case in C3 doesn't exist.
  • I haven't read the text of the GPL or the PSF survey data directly.
  • C6's "derivative work" boundary is a legal question that needs a lawyer.

Limitations

  • This tool accelerates human verification; it does not replace it.
  • Web search results may not include the most recent information or paywalled sources, and none were used here.
  • The adversarial review uses the same underlying model that may have produced the original output. It catches many issues but cannot catch all of them.
  • A claim rated VERIFIED means a supporting source was found, not that the claim is definitely correct. Sources can be wrong too.
  • Claims rated PLAUSIBLE may still be wrong. The absence of contradicting evidence is not proof of accuracy.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so nothing was fetched from the web and no file was written. The example is fictional. No web search was enabled, so the model says it could not run the live source checks and rated no claim as verified.

できること

AI の出力を 3 層で検証します。第 1 層では主張(事実、統計、引用、エンティティ、因果、時間)を抜き出し、内部の矛盾を確認します。第 2 層では Web 検索で裏づけまたは反証となる出典を探し、リンクを示します。第 3 層では、捏造された引用、出典のない正確すぎる数字、もっともらしいが誤った関連づけ、条件の抜け落ちといった幻覚パターンのチェックリストに沿って、敵対的にレビューします。各主張は VERIFIED、PLAUSIBLE、UNVERIFIED、DISPUTED、FABRICATION RISK のいずれかで評価されます。

動き方

  1. 単発モードでは、文章または以前の出力を指定すると、完全なレポートが返ります。
  2. 常時モードでは、会話の間ずっとオンのままで、内容のある回答に短い検証欄を付け、疑わしい主張や法律・規制に関する内容があれば完全なレポートに切り替わります。
  3. いつでも指示してオフにできます。

向いている場面

要約、調査メモ、判例・統計・規格を引用した下書きを、信頼する前に確認したいとき。

補足とリスク

指示のみのパッケージで、スクリプトはなく、ファイルも書き込みません。第 2 層にはエージェント側の Web 検索ツールが必要で、文章中の主張が検索クエリになるため、機密性のある文章は検証しないでください。Web に接続できない場合、この Skill は実際の確認ができなかったと述べます。目的は結論ではなく出典を示すことで、VERIFIED は裏づけとなる出典が見つかったという意味にすぎず、敵対的レビューも同種のモデルが行います。弁護士、医師などの専門家の代わりにはなりません。常時モードは、オフにするまで続きます。