0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

ClaudeのSSE受信、チャンク単位のdecodeで日本語が壊れる

0
Posted at

ClaudeのSSEを受けるなら、断片ごとのdecode()は避けたい。日本語の途中で切れると例外になり、errors="replace"で通すと文字が壊れる。受信したバイトをため、続きが届くまで待つ必要がある。

今日は「猫」の1バイト目の後で切った。置換も試したら、JSONの解析は通ったのに、本文が「���」になった。これは困るんだよね。

ClaudeのSSEと、受信チャンクの区切り

SSEはevent:でイベント名、data:でデータを書き、空行で区切る。HTTPの受信チャンクは、その空行やJSONの終わりに合わせて切れるとは限らない。UTF-8の「猫」は3バイト。その途中も試験に含めたい。

content_block_deltaにtext_deltaを入れた人工のSSEデータで、バイトを読み戻す部分を試す。Claude APIへの通信はしていない。

公式Python SDKのanthropic 0.102.0を手元で読んだ。SSEDecoderの実装は、バイトをためて空行を見つけ、各行をUTF-8でデコードする。

日本語の途中で切って、同じデータを読む

Python 3.12.13、anthropic==0.102.0で実行済み。sse_split.pyに保存し、同じ環境のPythonで実行する。APIキーは不要だ。

SSEDecoderはSDK内部のクラス。バージョンを固定した検証用に使う。本番コードから直接使う用途には向かない。

import json
from anthropic._streaming import SSEDecoder

wire = (
    'event: content_block_delta\n'
    'data: {"type":"content_block_delta","index":0,'
    '"delta":{"type":"text_delta","text":"猫"}}\n\n'
).encode('utf-8')
cut = wire.index('猫'.encode('utf-8')) + 1
parts = [wire[:cut], wire[cut:]]
try:
    ''.join(p.decode('utf-8') for p in parts)
except UnicodeDecodeError as exc:
    print(f'{type(exc).__name__}: {exc}')

def read_text(chunks):
    event = next(SSEDecoder().iter_bytes(iter(chunks)))
    return event.json()['delta']['text']

broken = ''.join(p.decode('utf-8', errors='replace') for p in parts)
print('replace:', json.loads(broken.split('data: ', 1)[1])['delta']['text'])
print('SDK:', read_text(parts))
for pos in range(1, len(wire)):
    assert read_text([wire[:pos], wire[pos:]]) == '猫'
assert read_text([wire[i:i+1] for i in range(len(wire))]) == '猫'
print(f'全{len(wire) - 1}か所の2分割と1バイト分割: OK')

出力はこうなった。

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe7 in position 110: unexpected end of data
replace: ���
SDK: 猫
全117か所の2分割と1バイト分割: OK

置換後もJSONは通る。

replaceは読めなかったバイトを置換文字に変える。後で文字列をつないでも、「猫」には戻せない。回答の保存処理なら、壊れた文字が記録に残る。

decodeを直した後にも、SSEの区切りがある

受信チャンクごとのjson.loads()も、途中で切れたJSONを解析してしまう。UTF-8の復元と、イベント単位への組み立ては両方必要だ。

このSDKでは_iter_chunks()が未完のバイトを保持し、iter_bytes()が行をデコードする。続くdecode()はdata:の行を集め、空行で一つのイベントにする。複数のdata:行も改行でつなぐ作りだった。

試験は118バイトの1イベントだけ。117か所は二つに切れる全位置の数だ。1バイト分割も通ったが、接続切断やエラーイベント、長い応答は未検証。

受信処理のテストに残すもの

自分なら、日本語を含むSSEの1バイト分割をテストに残す。英語だけだと今回の例外は出ない。Claude受信はSDKの公開APIに任せ、独自の中継では文字の復元とイベントの組み立てを確認する。SSE中継が増えるほど、この短い入力の出番も増えそうだ。

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?