0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

GPT-Redとは何か。OpenAIの自動red-teamingの仕組み・数値・限界を整理

0
Posted at

本記事は筆者が運営する AI Quotidia の海外ニュース解説記事です。

OpenAIは現地2026年7月15日、自社モデルを攻撃して弱点を洗い出す内部専用のLLM「GPT-Red」を発表した (MIT Technology Review独占報道)。

要点

  • 防御側モデルとのself-playループで攻撃を自己学習し、Webブラウジング/メール・カレンダー閲覧/コード編集を模した「dojo」環境で訓練された
  • 新型攻撃「fake chain of thought」を発見: 相手モデルの内部推論ログに偽エントリを挿入し、偽情報を「検証済み」として行動させる
  • 数値は3系統を区別する: ①最強攻撃群の成功率がGPT-5 (2025年8月) に90%超→GPT-5.6に23%未満 (MIT TR)。②間接プロンプトインジェクションのarena再現でGPT-Red 84% vs 人間red-teamer 13% (OpenAI公式ブログのリレー)。③fake chain of thought攻撃限定でGPT-5.1に約95%→GPT-5.6 Solに10%未満 (同)。対象攻撃もモデルも行ごとに別で、合成不可。いずれもOpenAIの自己申告で、第三者検証はまだない

完全版では「エージェント導入を進める日本の企業にとっての意味」の解説と FAQ、Quotidia の視点 (全文) を掲載しています。続きを読む

一次ソース

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?