Claude Sonnet 5.5が出ましたね!当たり前のようにAWSのAmazon Bedrockでも使えます。
ドキュメントもすでに更新されています。(仕事が早い!)
「RuntimeエンドポイントとMantleエンドポイントに対応、おすすめはRuntimeよ」ってのもいつも通りですね。
で、その下に見慣れないモノが!
Explicit Prompt Caching(明示的なプロンプトキャッシュ)は前からありましたが、Implicit Prompt Caching(暗黙的なプロンプトキャッシュ)は知らないデス!
確かOpenAIのGPT系は何も指定しなくてもキャッシュされた気がする。Claudeにもこの暗黙的が来たのですか??
Anthropic APIにも無い気がします!
やる!!
やってみた
プロジェクトを作って、
uv init app --python 3.13
cd $_
Boto3入れて、
uv add boto3
(Claude Codeが)スクリプト書いて、
import time
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
model_id = "global.anthropic.claude-sonnet-5-5"
# キャッシュの最小トークン数を超えるよう、長いシステムプロンプトを用意
system_prompt = "You are a helpful assistant.\n" + "\n".join(
f"Rule {i}: When the user asks about topic #{i}, answer politely and cite section {i}."
for i in range(1, 251)
)
for i in range(3):
response = client.converse(
modelId=model_id,
system=[{"text": system_prompt}], # cachePoint は付けない
messages=[{"role": "user", "content": [{"text": "Reply with just 'ok'."}]}],
inferenceConfig={"maxTokens": 5},
)
print(response["usage"])
time.sleep(3)
実行!
uv run main.py
結果
{'inputTokens': 7280, 'outputTokens': 4, 'totalTokens': 7284, 'cacheReadInputTokens': 0, 'cacheWriteInputTokens': 0}
{'inputTokens': 7280, 'outputTokens': 4, 'totalTokens': 7284, 'cacheReadInputTokens': 0, 'cacheWriteInputTokens': 0}
{'inputTokens': 7280, 'outputTokens': 4, 'totalTokens': 7284, 'cacheReadInputTokens': 0, 'cacheWriteInputTokens': 0}
キャッシュ効いとらんやないか!
Anthropic APIの方法を試してみる。
Anthropic APIにはAutomatic caching(自動キャッシング)という仕組みがあります。
こんな感じでcache_controlパラメーターを指定します。(参考)
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
cache_control={"type": "ephemeral"},
system="You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.",
messages=[
{
"role": "user",
"content": "Analyze the major themes in 'Pride and Prejudice'.",
}
],
)
print(response.usage.model_dump_json())
Automatic cachingと対になる方法としてはExplicit cache breakpoints(明示的なキャッシュブレークポイント)という物があり、これは、cache_controlをメッセージ、システムプロンプト、ツール中に含める必要があります。
で、この方法を、BedrockのConverse APIで試してみます。昔やったときはうまくいきませんでしたが。。
import time
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
model_id = "global.anthropic.claude-sonnet-5-5"
system_prompt = "You are a helpful assistant.\n" + "\n".join(
f"Rule {i}: When the user asks about topic #{i}, answer politely and cite section {i}."
for i in range(1, 251)
)
for i in range(3):
response = client.converse(
modelId=model_id,
system=[{"text": system_prompt}],
messages=[{"role": "user", "content": [{"text": "Reply with just 'ok'."}]}],
inferenceConfig={"maxTokens": 5},
# Anthropic API の Automatic caching と同じく、最上位に cache_control を渡す
additionalModelRequestFields={"cache_control": {"type": "ephemeral"}},
)
print(response["usage"])
time.sleep(3)
additionalModelRequestFields={"cache_control": {"type": "ephemeral"}},の部分です
{'inputTokens': 4, 'outputTokens': 4, 'totalTokens': 7284, 'cacheReadInputTokens': 0, 'cacheWriteInputTokens': 7276, 'cacheDetails': [{'ttl': '5m', 'inputTokens': 7276}]}
{'inputTokens': 4, 'outputTokens': 4, 'totalTokens': 7284, 'cacheReadInputTokens': 7276, 'cacheWriteInputTokens': 0}
{'inputTokens': 4, 'outputTokens': 4, 'totalTokens': 7284, 'cacheReadInputTokens': 7276, 'cacheWriteInputTokens': 0}
キャッシュ効く!!
そしてなんと、Sonnet 5.5以外のモデルも効く!!
キャッシュが有効になったモデル
Opus 5.5、Sonnet 5.5、Opus 5、Sonnet 5、Opus 4.8、Opus 4.7、Haiku 4.5
Sonnet 4.6、Opus 4.6はキャッシュ効かなかったー
まとめ
思ってた暗黙的ではありませんが、簡単プロンプトキャッシュは実装されました!めでたい!

