timeout=0.1なら、ツールも0.1秒で止まる。そう思ってエージェントの締切に使うと危ない。手元で0.6秒かかる処理を走らせたら、Future.result(timeout=0.1)は例外を返したのに、呼び出し側は約0.7秒待った。子プロセスに分けて停止させた場合は約0.1〜0.2秒だった。
自分はツールの制限時間を付けるとき、「待つのをやめる」と「実行を止める」を分けて考える。前者だけでは処理が走り続け、次の実行と重なる。
Q. Future.result(timeout)で何が起きる?
次のコードは標準ライブラリだけで動く。Macで0.6秒眠る関数と子プロセスを同じ0.1秒で待った。計測範囲は起動前から片付け終わるまで。
from concurrent.futures import ThreadPoolExecutor, TimeoutError as FutureTimeout
import os
import signal
import subprocess
import sys
import time
def slow_tool():
time.sleep(0.6)
return "done"
start = time.monotonic()
with ThreadPoolExecutor(max_workers=1) as pool:
future = pool.submit(slow_tool)
try:
future.result(timeout=0.1)
except FutureTimeout:
print("thread: timeout, cancel=", future.cancel())
print(f"thread: {time.monotonic() - start:.2f}s")
start = time.monotonic()
proc = subprocess.Popen(
[sys.executable, "-c", "import time; time.sleep(0.6)"],
start_new_session=True,
)
try:
proc.wait(timeout=0.1)
except subprocess.TimeoutExpired:
os.killpg(proc.pid, signal.SIGKILL)
proc.wait()
print("process: killed")
print(f"process: {time.monotonic() - start:.2f}s")
最初の実行では、こう出た。
thread: timeout, cancel= False
thread: 0.73s
process: killed
process: 0.11s
3回続けると、スレッド側は0.727〜0.749秒、子プロセス側は0.117〜0.217秒だった。環境を問わず0.2秒以内に終わる保証はない。タイトルの数字は観測値を小数1桁に丸めたもの。
Q. 例外が出たのに、なぜ0.7秒待つ?
future.result(timeout=0.1)が打ち切るのは、結果を待つ時間。実行中のslow_tool()には停止指示が届かない。走り始めたFutureへのcancel()はFalseを返した。with ThreadPoolExecutor(...)を抜けるときも、既定のshutdown(wait=True)で完了を待つ。0.1秒で例外を見た後、残りを待っていた。
shutdown(wait=False)に替えてもスレッドは残る。ツールがファイルを書いたり外部要求を送ったりすれば、応答後に副作用が出る。cancel=Falseを見たとき、ここがいちばん気になった。
同じ0.1秒の指定でも、止めている対象が違う。
子プロセス側は時間切れ後にSIGKILLを送り、wait()で回収する。start_new_session=Trueで専用のプロセスグループを作り、さらに子を起動したときも停止対象に入れた。Popenには引数を配列で渡す。この部分はPOSIX向けで、Windowsにはそのまま移せない。
Q. ツールが2回続くとき、締切はどう渡す?
各ツールにtimeout=1.0を渡すと、2回で最大約2秒待つ。応答全体の締切が1秒なら、開始時刻から絶対時刻を一つ決め、各ツールの直前に残り時間を計算する。
import os
import signal
import subprocess
import sys
import time
def run_until(argv, deadline):
remaining = deadline - time.monotonic()
if remaining <= 0:
raise TimeoutError("deadline passed before launch")
proc = subprocess.Popen(argv, start_new_session=True)
try:
return proc.wait(timeout=remaining)
except subprocess.TimeoutExpired:
os.killpg(proc.pid, signal.SIGKILL)
proc.wait()
raise TimeoutError("tool deadline exceeded") from None
start = time.monotonic()
deadline = start + 1.0
for delay in (0.08, 2.0):
try:
run_until(
[sys.executable, "-c", f"import time; time.sleep({delay})"],
deadline,
)
print(f"{delay}: done")
except TimeoutError as exc:
print(f"{delay}: {exc}")
print(f"total: {time.monotonic() - start:.2f}s")
手元の結果。
0.08: done
2.0: tool deadline exceeded
total: 1.04s
1秒をぴったり守れてはいない。子プロセスの起動、シグナル送信、回収にも時間がかかる。外向きの応答に厳密な締切があるなら、その分を見込んでツール用の締切を短く置く。
Q. 実務のツールへそのまま当てていい?
止められるのは、自分で起動した外部コマンドだ。SDKやHTTPクライアントには、その呼び出しに合った接続・読み取りの制限時間も設定する。外部コマンドへ分離するなら、入力と出力の受け渡しや終了後のファイル処理も要る。
SIGKILLは片付け処理を実行させない。途中まで書いたファイルは残り、送信済みの外部要求も取り消せない。試したコードは0.6秒の待機を止める検証だ。実処理では出力先を一時領域へ分け、成功した結果だけを確定させたい。
タイムアウトの例外だけで安心すると、待機と実行の差を見落とす。外部コマンド型のツールなら、終了シグナルを送り、回収まで測る。複数ツールには共通の絶対締切を渡す。締切を守れたかはresult(timeout=...)の値ではなく、呼び出し元が実際に戻った時刻で確認する。