Haiku 5.5へ移行するときは、モデルIDを交換する前に、応答を受け取る処理と料金見積りの前提を分離しておきたいです。この記事では、思考ブロックを含む応答、途中で終わった応答、使用量の欠損を検査する小さなPythonコードを作ります。
扱うのは2026年10月8日に確認したClaude APIの公式仕様と、ネットワークを使わない合成fixtureです。HaikuへのAPI送信、実モデルの推論、課金実験は行っていません。最後まで記事だけで再現できる構成にします。
移行で変わる三つの前提
正式モデルIDは claude-haiku-5-5。最大100万tokenのコンテキストと、通常の最大出力12万8000tokenがあります。ただし、入力が10万token以下なら入力0.10ドル・出力0.50ドル、10万tokenを超える場合は0.50ドル・2.50ドルという基本単価の区分があります。いずれも100万tokenあたりです。文脈の容量を、そのまま安い単価の適用範囲と考えないようにします。
移行ガイドには、同じテキストでもHaiku 4.5よりtoken数が概ね30%増えるとの説明があります。これは文字数から必ず1.3倍に換算できるという仕様でも、この記事でtoken数を測った結果でもありません。入力内容によって変わります。
さらに、adaptive thinkingが既定で、effortはmediumが既定です。以前の budget_tokens、標準外のサンプリング値など、エラーになる設定もあるため、既存の設定オブジェクトを丸ごと引き継ぐ前に確認します。
資料:Claude API:Haiku 5.5の仕様、Claude API:移行ガイド
先頭ブロックを回答と決めつけない
content[0]["text"] だけを読む実装は、先頭が思考ブロックだったときに壊れます。ブロックの type を見て、表示するテキストと、別処理が必要なブロックを区別します。
また、文字列が取れたことと、回答が完了したことは別です。出力上限で終了した max_tokens、ツール実行へ移る tool_use などを、完了したテキストとして登録すると、途中の内容を確定結果にしてしまいます。
下の例は、最終JSONの end_turn と表示テキストを扱うための限定した処理です。ツール実行、ストリーミング中の差分、途中再開は実装しません。未知のブロックや終了理由を受けたときは、例外を返して用途別の処理へ渡す設計です。
usageの欠損を0で埋めない
キャッシュを使う入力は、新規入力、キャッシュ作成、キャッシュ読取のカウンターに分かれます。新規入力だけを見ると、キャッシュした大きな資料を見落とす可能性があります。思考も課金対象の出力に含まれるため、画面に表示された本文の文字数だけで出力料金を見積もりません。
この例では、正規化した使用量として次の4項目を要求します。欠損値を get(key, 0) で補わず、非負の整数であることを確かめます。Pythonではboolがintの派生なので、isinstance(value, int) だけで True を1tokenとして受け付けないようにします。
input_tokenscache_creation_input_tokenscache_read_input_tokensoutput_tokens
本番のアダプターでは、使用量が提供される条件、集計範囲、キャッシュの内訳を公式仕様に合わせて正規化する必要があります。ここでは合成データをこの形式で作ります。
資料:Claude API:コンテキストと使用量、Claude API:プロンプトキャッシュ
料金表と、料金計数の実装を分ける
料金表の単価は確認できても、見積りの実装には、どの数をプロンプトの境界判定へ使うか、キャッシュをどう扱うか、料金帯をリクエスト全体へどう適用するかという契約が必要です。この記事では、この詳細を実アカウントの請求と対応づけて確認していません。
そのため、以下の FixtureBand は本番の課金判断を自動生成するクラスではありません。合成fixtureが仮定したプロンプト数と、使用量のハッシュを明示して持つための値です。実usageの合計を、未確認の料金条件へ自動変換しません。
計算例は「このfixtureでは、明示した入力の帯の単価をリクエスト全体へ適用する」と仮定しています。5分キャッシュ作成と読取の単価も使いますが、1時間キャッシュ、バッチ割引、ツール料金、特典、税、通貨換算は対象外です。本番の見積り関数は停止するままにしてあり、これらの根拠を確認してから別アダプターとして実装します。
資料:Claude API:Haiku 5.5の基本・キャッシュ単価、Claude API:料金表
最小コード
Python 3.9以上の標準ライブラリだけで動作します。SDK、APIキー、モデルのダウンロードは不要です。次を haiku_migration_guard.py として保存します。
"""Offline synthetic fixture guard. No network, SDK, API key or billing calls."""
from dataclasses import dataclass
from decimal import Decimal
import hashlib
import json
MODEL = "claude-haiku-5-5"
USAGE_KEYS = ("input_tokens", "cache_creation_input_tokens",
"cache_read_input_tokens", "output_tokens")
class Reject(ValueError):
pass
def count(value, name):
if type(value) is not int or value < 0:
raise Reject(f"{name}: non-negative integer required")
return value
def checked_usage(response):
usage = response.get("usage")
if not isinstance(usage, dict):
raise Reject("usage missing")
return {key: count(usage.get(key), key) for key in USAGE_KEYS}
def usage_digest(usage):
raw = json.dumps(usage, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(raw.encode()).hexdigest()
def complete_text(response):
if not isinstance(response, dict) or response.get("model") != MODEL:
raise Reject("unexpected response model/shape")
if response.get("stop_reason") != "end_turn":
raise Reject("response is not a confirmed complete text turn")
content = response.get("content")
if not isinstance(content, list) or not content:
raise Reject("content missing")
texts = []
for block in content:
if not isinstance(block, dict):
raise Reject("invalid block")
kind = block.get("type")
if kind == "text":
if not isinstance(block.get("text"), str):
raise Reject("text field missing")
texts.append(block["text"])
elif kind in ("thinking", "redacted_thinking"):
continue
else:
raise Reject(f"unhandled block: {kind}")
text = "".join(texts)
if not text.strip():
raise Reject("no visible text")
return text
@dataclass(frozen=True)
class FixtureBand:
prompt_tokens: int
usage_sha256: str
scope: str = "synthetic-fixture-request-wide-band-assumption"
def fixture_estimate(response, band):
"""Prices an explicit fixture assumption, never an actual invoice."""
text = complete_text(response)
usage = checked_usage(response)
if not isinstance(band, FixtureBand):
raise Reject("prompt band evidence missing; do not infer it")
if band.scope != "synthetic-fixture-request-wide-band-assumption":
raise Reject("live prompt/cache tariff accounting is not verified here")
prompt = count(band.prompt_tokens, "prompt_tokens")
if usage_digest(usage) != band.usage_sha256:
raise Reject("band evidence belongs to different usage")
if prompt > 1_000_000:
raise Reject("fixture exceeds model context capacity")
# Request-wide band application is a fixture assumption, not a live claim.
rates = ("0.10", "0.125", "0.01", "0.50") if prompt <= 100_000 \
else ("0.50", "0.625", "0.05", "2.50")
usd = sum(Decimal(usage[key]) * Decimal(rate)
for key, rate in zip(USAGE_KEYS, rates)) / Decimal(1_000_000)
return {"text": text, "fixture_usd": str(usd), "scope": band.scope}
def forbid_unverified_live_estimate(response):
"""Keep actual tariff adaptation closed until primary/account proof exists."""
complete_text(response)
checked_usage(response)
raise Reject("live estimate disabled: verify prompt counting and band application")
complete_text() は先頭に思考ブロックがあっても、typeがtextの部分を収集します。thinkingやredacted_thinkingは表示テキストへ混ぜません。未知のtypeを無視して処理を続ける代わりに、未対応として停止します。
fixture_estimate() は使用量と FixtureBand のハッシュが一致しない場合にも停止します。使用量を変更したのに以前の境界判断だけを再利用することを防ぐためです。ただし、このハッシュは料金条件の正しさを証明しません。合成データ内の対応関係を守るもので、本番の根拠確認を置き換える機能ではありません。
手元で確認する
同じフォルダで、次の最小例を実行できます。
from haiku_migration_guard import (
MODEL, FixtureBand, checked_usage, fixture_estimate, usage_digest,
)
sample = {
"model": MODEL,
"stop_reason": "end_turn",
"content": [
{"type": "thinking", "thinking": "合成の思考ブロック"},
{"type": "text", "text": "決定事項:資料を確認する"},
],
"usage": {
"input_tokens": 1000,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0,
"output_tokens": 200,
},
}
fixture = FixtureBand(1000, usage_digest(checked_usage(sample)))
print(fixture_estimate(sample, fixture))
結果の fixture_usd は 0.0002、textは「決定事項:資料を確認する」となります。200tokenは課金対象の出力をまとめた仮定で、実モデルの出力や実請求ではありません。
19ケースを再現するには、次を run_fixtures.py として先ほどのコードと同じフォルダへ保存します。認証情報や追加ライブラリは不要です。
"""Meaningful failure fixtures; no SDK/model/API execution."""
from copy import deepcopy
import json
import platform
import sys
from haiku_migration_guard import (MODEL, FixtureBand, Reject, checked_usage,
fixture_estimate, forbid_unverified_live_estimate, usage_digest)
def response(input_tokens=1000, output_tokens=200):
return {"model": MODEL, "stop_reason": "end_turn", "content": [
{"type": "thinking", "thinking": "Synthetic, not model output"},
{"type": "text", "text": "決定事項:資料を確認する"}],
"usage": {"input_tokens": input_tokens, "cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0, "output_tokens": output_tokens}}
def band(r, tokens):
return FixtureBand(tokens, usage_digest(checked_usage(r)))
def main():
results = []
for name, r, tokens, expected in [
("basic_arithmetic", response(), 1000, "0.0002"),
("boundary_100000", response(100000, 200), 100000, "0.0101"),
("boundary_100001", response(100001, 200), 100001, "0.0505005")]:
value = fixture_estimate(r, band(r, tokens))
assert value["fixture_usd"] == expected
assert value["text"] == "決定事項:資料を確認する"
results.append({"case": name, "accepted": True, "result": value})
cached = response(50, 200)
cached["usage"]["cache_read_input_tokens"] = 100000
# The 100050 count is authored as a synthetic assumption, not API inference.
value = fixture_estimate(cached, band(cached, 100050))
assert value["fixture_usd"] == "0.005525"
results.append({"case": "explicit_cached_fixture", "accepted": True,
"result": value})
base = response(); old_band = band(base, 1000)
failures = []
for name, mutate in [
("missing_usage", lambda r: r.pop("usage")),
("missing_cache_count", lambda r: r["usage"].pop("cache_read_input_tokens")),
("negative_count", lambda r: r["usage"].update(input_tokens=-1)),
("boolean_count", lambda r: r["usage"].update(output_tokens=True)),
("truncated_turn", lambda r: r.update(stop_reason="max_tokens")),
("tool_turn", lambda r: r.update(stop_reason="tool_use")),
("unhandled_block", lambda r: r["content"].append({"type":"tool_use"})),
("unknown_block", lambda r: r["content"].append({"type":"future_type"})),
("text_field_missing", lambda r: r["content"][1].pop("text")),
("thinking_only", lambda r: r["content"].pop()),
("model_mismatch", lambda r: r.update(model="other-model")),
("stale_usage_evidence", lambda r: r["usage"].update(input_tokens=1001))]:
r = deepcopy(base); mutate(r)
failures.append((name, lambda r=r: fixture_estimate(r, old_band)))
failures += [
("missing_band", lambda: fixture_estimate(base, None)),
("unverified_live_band", lambda: fixture_estimate(base,
FixtureBand(1000, old_band.usage_sha256, "live-usage"))),
("actual_live_estimate_disabled", lambda: forbid_unverified_live_estimate(base))]
for name, call in failures:
try:
call()
except Reject as e:
results.append({"case":name,"rejected":True,"reason":str(e)})
else:
raise AssertionError(f"failure fixture accepted: {name}")
print(json.dumps({"status":"pass", "python":sys.version,
"platform":platform.platform(), "positiveCases":4, "negativeCases":15,
"networkOrPaidApiCalls":0, "HaikuInferencePerformed":False,
"cases":results}, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
そのフォルダで、次を実行します。
python3 run_fixtures.py
出力JSONには status: "pass"、positiveCases: 4、negativeCases: 15 と各ケースの結果が表示されます。予期した拒否が起きない場合や演算値が異なる場合は、スクリプトがエラーで終了します。
このfixture一式で、次を確認しました。
- 入力1000・出力200の基本演算:0.0002ドル。
- 入力10万・出力200のfixture:0.0101ドル。
- 入力10万1・出力200のfixture:0.0505005ドル。
- キャッシュ読取10万・新規入力50・出力200で、境界入力を10万50と明示したfixture:0.005525ドル。
後二つの数値は、上で述べたリクエスト全体への料金帯適用というfixtureの仮定を含みます。実APIの請求として使う数値ではありません。
正常4ケースに加えて、使用量の欠損、キャッシュカウンターの欠損、負数、bool、途中終了、ツールturn、未対応ブロック、文字列欠損、思考だけの応答、別モデル、使用量変更後の古い根拠、境界根拠の欠損、未確認のlive指定など15ケースを拒否することを確認しました。本番見積り関数も、条件未確認として停止します。
実行環境は2026年10月8日のmacOS 27.0.1 arm64、Python 3.9.6。ネットワーク呼出し0、SDK使用0、Haiku推論0です。確認したのは合成データに対する検査とDecimalの演算で、Haikuの精度、速度、請求額ではありません。
本番へ接続するときに残る作業
最初に、使うAPIとアカウント条件に対応した入力計数・料金帯・キャッシュの仕様を確定し、取得元と確認日を料金アダプターへ記録します。未確認なら見積りを止める現在の処理を維持します。使用量が欠けたときに「無料だった」と解釈しないようにします。
次に、end_turn以外の終了理由、ツール利用、ストリーミング再開を、アプリの状態遷移として実装します。途中で得た文字列は確定データに登録せず、完了・未完了を保存する形にします。
最後に、同じ入力を少数だけ実行して、出力、使用量、待ち時間、料金条件を照合します。これは今後の確認手順の提案です。この記事でそのAPI試験を実施したという意味ではありません。
モデルIDだけを交換する移行よりも、応答の型と終了状態、使用量、料金条件を別々に検査できる境界を先に作ると、仕様が変わったときに確認する場所を絞れます。