2
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

PR: データブリックス・ジャパン株式会社
Omnigentによるメタハーネス入門(1)基礎編

MLflow 3.14の @mlflow.test で、エージェントの評価をpytestのテストにする

2
Posted at

はじめに

2026年6月にリリースされたMLflow 3.14で、@mlflow.test というデコレータが追加されました。

mlflow.genai.evaluate() でエージェントを評価した結果を、そのままpytestのテストとして扱える機能です。スコアラーやLLMジャッジはこれまでと同じものを使い、評価の結果を assert で判定するだけでテストになります。

プロンプトを変えた、モデルを差し替えた、といった変更のたびに評価を回している方は多いと思います。その評価をpytestのテストとして書いておけば、CIで実行して、品質が落ちた変更をそこで止められます。

Databricks Free Editionで、合格するテストと不合格になるテストを実際に流してみました。結果がMLflowにどう残るかも含めて書いていきます。

@mlflow.test とは

普段のpytestのテスト関数に @mlflow.test を付け、中で mlflow.genai.evaluate() を呼び、その結果を assert します。

@mlflow.test
def test_answers_concisely(agent):
    result = mlflow.genai.evaluate(
        predict_fn=agent,
        data=[{"inputs": {"question": "What are your hours?"}}],
        scorers=[Guidelines(name="concise", guidelines="Answer in one sentence.")],
    )
    assert result.passed, result.reason

ポイントは evaluate() の戻り値に付いている2つの属性です。

属性 中身
result.passed すべてのスコアラーが、すべての行で合格したときだけ True
result.reason 不合格になったスコアラーの名前と、その理由

評価の部分は何も変わりません。新しく覚えるのは、デコレータと assert result.passed, result.reason の1行です。

テストとして書くので、pytestの機能はそのまま使えます。parametrize で複数の質問を流すこともできますし、CIでは普通に pytest を実行するだけです。

試した環境

  • Databricks Free Edition (サーバーレスのノートブック)
  • MLflow 3.16.1
  • pytest 9.1.1

テストを書く

ライブラリのインストール

@mlflow.test はMLflow 3.14以降で使えます。

%pip install -U "mlflow[databricks]>=3.14" pytest
dbutils.library.restartPython()

実験を設定する

評価の結果を記録する実験を設定します。

import os
import sys
import mlflow
import pytest

USER = spark.sql("SELECT current_user()").first()[0]
experiment = mlflow.set_experiment(f"/Users/{USER}/mlflow_test_demo")
# このあと pytest をノートブックと同じプロセスで動かすので、環境変数にも入れておきます
os.environ["MLFLOW_EXPERIMENT_ID"] = experiment.experiment_id

TEST_DIR = "/tmp/mlflow_test_demo"
os.makedirs(TEST_DIR, exist_ok=True)

エージェントを用意する

サポート窓口のエージェントを想定します。今回は仕組みを見ることが目的なので、エージェントは決まった文を返すだけの関数にしました。

  • good_agent: 英語で1文だけ答える
  • bad_agent: 日本語で2文答える

実際に使うときは、ここを自分のエージェントに置き換えます。pytestのフィクスチャとして conftest.py に書いておきます。

CONFTEST = '''
import pytest


@pytest.fixture
def good_agent():
    def agent(question: str) -> str:
        return "Our support hours are 9am to 5pm on weekdays."
    return agent


@pytest.fixture
def bad_agent():
    def agent(question: str) -> str:
        return "サポート時間は平日9時から17時です。土日祝日はお休みです。"
    return agent
'''

評価の結果を assert する

テストでは「英語で答えること」「1文で答えること」の2つを Guidelines (LLMジャッジ) で評価します。good_agent のテストは parametrize で質問を2つ流し、bad_agent のテストは質問1つです。

TEST_SUPPORT_AGENT = '''
import mlflow
import pytest
from mlflow.genai.scorers import Guidelines

GUIDELINES = [
    Guidelines(name="is_english", guidelines="The answer must be written in English."),
    Guidelines(name="is_concise", guidelines="The answer must be a single sentence."),
]

QUESTIONS = [
    "What are your support hours?",
    "When can I contact support?",
]


@mlflow.test
@pytest.mark.parametrize("question", QUESTIONS)
def test_good_agent(good_agent, question):
    result = mlflow.genai.evaluate(
        predict_fn=good_agent,
        data=[{"inputs": {"question": question}}],
        scorers=GUIDELINES,
    )
    # すべてのスコアラーが合格したときだけ passed が True になります
    assert result.passed, result.reason


@mlflow.test
def test_bad_agent(bad_agent):
    result = mlflow.genai.evaluate(
        predict_fn=bad_agent,
        data=[{"inputs": {"question": "What are your support hours?"}}],
        scorers=GUIDELINES,
    )
    assert result.passed, result.reason
'''

for name, body in [("conftest.py", CONFTEST), ("test_support_agent.py", TEST_SUPPORT_AGENT)]:
    with open(os.path.join(TEST_DIR, name), "w") as f:
        f.write(body.lstrip())

evaluate() の呼び方は、普段の評価とまったく同じです。違いはデコレータと最後の assert だけです。

pytest で実行する

MLflowのpytestプラグインを読み込む

-p mlflow.pytest.plugin でMLflowのpytestプラグインを読み込みます。プラグインを読み込むと、1回のpytest実行で動いたテストの評価が、1つのMLflowランにまとまります。

ここが最初のハマりどころです。プラグインを読み込まなくてもテストの合否は判定されますが、evaluate() を呼ぶたびに別々のランが作られました。デコレータを付けたからといって、ランがまとまるわけではありません。

普段は pyproject.toml に書いておき、ターミナルで pytest を実行します。

[tool.pytest.ini_options]
addopts = ["-p", "mlflow.pytest.plugin"]

今回はノートブックから pytest.main() で実行しました。

# 同じノートブックで再実行できるよう、前回読み込んだテストモジュールを捨てます
for mod in [m for m in sys.modules if m == "conftest" or m.startswith("test_")]:
    del sys.modules[mod]

exit_code = pytest.main([
    TEST_DIR,
    "-v",
    "--import-mode=importlib",
    "-p", "no:cacheprovider",
    # MLflowのpytestプラグイン。これがないとテストごとに別々のランになります
    "-p", "mlflow.pytest.plugin",
])
print("終了コード:", int(exit_code))

合格と不合格の出方

実行すると、good_agent の2件が合格、bad_agent が不合格になりました。

test_support_agent.py::test_good_agent[What are your support hours?] PASSED [ 33%]
test_support_agent.py::test_good_agent[When can I contact support?] PASSED [ 66%]
test_support_agent.py::test_bad_agent FAILED [100%]

行末の [ 33%] は評価のスコアではありません。3件のテストのうち何件目まで終わったかを示す、pytestの進み具合の表示です。

不合格になったテストには、result.reason の内容がそのまま表示されます。

>       assert result.passed, result.reason
E       AssertionError: 2 assertions failed:
E           - is_concise: The guideline states that the answer must be a single sentence. The provided response contains multiple sentences, which does not comply with this guideline. Therefore, the guideline is not satisfied.
E           - is_english: The guideline states that the answer must be written in English. However, the response provided is in Japanese. Therefore, the response does not comply with the guideline that requires the answer to be in English.

どのジャッジが、どういう理由で不合格にしたのかまで、pytestの出力だけで分かります。CIのログを見た時点で、何が壊れたのかの見当が付く。テストとして使ううえで、この点はかなり効いてきます。

テストが1件でも失敗すると、終了コードは1になります。

=================== 1 failed, 2 passed, 3 warnings in 12.69s ===================
終了コード: 1

CIでは、この終了コードでジョブが失敗し、品質が落ちた変更をそこで止められます。普段のユニットテストと同じ扱いです。

結果はMLflowの評価ランにまとまる

pytestの出力で分かるのは合否と理由までです。各テストでエージェントが何を返したのかは、MLflowのUIで確認します。

3件のテストの評価は、1つの評価ランにまとまっていました。評価ランの画面を開くと、テストごとのトレースと、それぞれに付いたジャッジの判定が一覧で見られます。

Screenshot 2026-10-04 at 20.35.20.JPG

上部には、全テストを通した合格率が出ます。今回は3件中2件の合格なので、is_concise と is_english のどちらも67%です。不合格の行を見れば、bad_agent が日本語で2文を返したことがそのまま分かります。

なお、トレースの「状態」列は3件ともOKになっています。この列はエージェントの処理が正常に終わったかを示すもので、テストの合否とは別物です。合否はジャッジの列で見ます。

ランの状態は、テストが1件でも失敗すると FAILED、すべて合格すると FINISHED でした。実験のラン一覧を見るだけで、どのpytest実行が失敗したのかが分かります。

一点だけ注意があります。mlflow.search_runs() で取得したランの指標 (is_english/mean など) には、最後に実行したテストの値が入っていました。今回は bad_agent の0.0です。全テストを通した集計は、評価ラン画面の合格率で見るのが確実です。

まとめ

@mlflow.test を試して分かったことをまとめます。

  • mlflow.genai.evaluate() の結果を assert result.passed, result.reason で判定すれば、pytestのテストになる
  • 評価の部分はこれまでと同じ。スコアラーやLLMジャッジをそのまま使える
  • 不合格になると、どのジャッジがどんな理由で不合格にしたかがpytestの出力に出る
  • 1件でも失敗すると終了コードが1になるので、CIでそのまま品質ゲートにできる
  • ランを1つにまとめるには、-p mlflow.pytest.plugin でプラグインを読み込む必要がある
  • テストごとのトレースとジャッジの判定は、MLflowの評価ラン画面で一覧できる

一番の収穫は、評価とテストが別々のものではなくなったことでした。これまでは評価を回して数字を眺め、良し悪しを人が判断していました。@mlflow.test を使えば、その判断をpytestに任せて、普段のテストと同じ流れでCIに載せられます。エージェントがおかしな応答を返した事例を見つけるたびに、テストとして1件ずつ足していく。そういう育て方が自然にできそうです。

参考リンク

はじめてのDatabricks

はじめてのDatabricks

Databricks無料トライアル

Databricks無料トライアル

2
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
2
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?