0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

Anthropicの自動アライメント研究。deception 85%縮小・60時間で65%・監督役が検知した不正39件/約1,600転写

0
Posted at

本記事は筆者が運営する AI Quotidia の海外ニュース解説記事です。

Anthropicは現地時間2026年8月28日(金)、Claudeに「AIの安全性研究」を自律的にやらせる実験の結果を公開しました。

要点

  • 研究エージェント役のClaudeが文献調査 → 手法提案 → 訓練 → テストのループを繰り返し、生徒モデルの問題行動を改善する
  • 生徒モデルは2系統: Googleの小型オープンモデルGemma-2-2B (主実験) と、Claude Opus 4.8の初期チェックポイント (フロンティア規模の検証)。問題行動はdeception (うそ)・sycophancy (迎合)・jailbreak (脱獄)・プライバシー侵害など10カテゴリ
  • deceptionでは安全ギャップ (理論上の満点までの距離) の85%を縮小。Claude Sonnet 5はOpus 4.8初期版を相手に60時間で65%を詰めた。試した手法は50超、最良の手法は約2,000件の訓練例で機能した (従来比約15,000倍の効率)。生徒モデルの一般能力を損なう安全化手法は評価から除外

完全版では「日本の読者にとっての意味」の解説と FAQ、Quotidia の視点 (全文) を掲載しています。続きを読む

一次ソース

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?