0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

Jevを試してみた

0
Posted at

背景・目的

Jevは、TypeSafe AIが2026年9月15日に公開した、System One Modelと呼ばれる新しいクラスのAIモデルです。

一般的なLLMが文章生成を主な出力とするのに対し、Jevはソフトウェアから直接利用できる、型付きの意思決定と確率・Confidenceを返すことに特化しているようです。

どのようなサービスなのか理解するため整理し、実際に試してみます。

まとめ

項目 内容
概要 Jevは、TypeSafe AIが公開した最初のSystem One Model。文章生成ではなく、ソフトウェア内で利用するDecisionの生成に特化
特徴 Type-safeな構造化されたDecisionとProbability / Confidenceを返す。複数の出力を並列に生成し、高速・低コストな処理を特徴とする
Decision Noul(Yes / No)、Choice(複数候補から選択)、Score(尺度上で評価)の3種類
ユースケース classify、route、score、extract、branchなど、自然言語や曖昧な状態を含む判断をソフトウェアのWorkflowに組み込む用途

概要

Jevとは

Jevは、TypeSafe AIが2026年9月15日に公開した、System One Model と呼ばれる新しいクラスのAIモデルです。TypeSafe AIでは、Jevを最初の公開System One Modelとして位置付けており、現在はEarly Accessとして提供しています。

Our first public model is Jev, available today in early access.

TypeSafe AIでは、Jevを文章生成ではなく、ソフトウェアから直接利用できる高速・構造化された意思決定に特化したモデルとして位置付けています。

System One Model

a new class of frontier models built to make fast, structured decisions that software can use directly.

System One Modelは、ソフトウェアから直接利用できる、高速かつ構造化された意思決定を行うためのモデルとして定義されています。

一般的なLLMが文章やコードなどの文字列を生成するのに対し、System One Modelは、アプリケーション内で利用するDecisionを返すことを目的としています。

名称は、Daniel Kahnemanの『Thinking, Fast and Slow』で示された、人間の高速で直感的な「System 1」と、低速で熟考する「System 2」の区別に由来しています。

Decisions, not strings

Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.

Jevは、非構造な状態(State)を入力し、型付けされた確率的なDecisionを出力する関数として捉えることができます。

LLMでは文字列を生成するため、ソフトウェアで利用する場合は出力のParseやValidationが必要になりますが、
一方、Jevでは出力候補や構造を事前に定義し、Type-safeな構造化された値としてDecisionを受け取ります。

State
  ↓
Jev
  ↓
Typed Decision
+ Probability / Confidence
  ↓
Application

Probability / Confidence

All answers are accompanied with calibrated probabilities and confidence scores.

Jevでは、Decisionだけでなく、確率(Probability)やConfidenceも合わせて返却されます。

TypeSafe AIでは、Confidenceが高いほど実際のAccuracyも高くなるように調整された「Calibrated Decisions」を特徴として挙げています。

これにより、例えばConfidenceに応じて、自動処理するか人による確認に回すかをアプリケーション側で制御できます。

Confidence 制御
高い 自動処理
低い Human Review

LLMとの違い

Outputs: Type-safe structured values. Possible outputs and structure are defined in advance.

一般的なLLMの出力は文字列であり、用途によってはParseやValidationが必要です。

Jevでは、取り得る出力と構造をあらかじめ定義し、その範囲内でDecisionを返します。また、TypeSafe AIはJevについて、型エラーが発生しないことを特徴として挙げています。

LLM Jev
主な出力 文字列 Type-safeなDecision
出力方法 Tokenを逐次生成 並列でDecisionを生成
ソフトウェア利用 Parse / Validationが必要 直接利用しやすい
不確実性 一貫したConfidence取得が難しい Probability / Confidenceを返す
主な用途 Chat、文章生成、Codingなど classify、route、score、branchなど

高速・低コスト

End-to-end response time is 70ms-500ms for TypeSafe.

Jevは、TypeSafe AIの公開情報ではエンドツーエンドで70〜500msの応答時間とされています。

System One Modelは、チャットのように人間との対話を前提とするのではなく、ソフトウェア内でリアルタイムに意思決定することを想定しているため、高速な応答を特徴としています。

Input tokens: $0.042 / MTok ($42 per billion tokens).
Output tokens: FREE (too cheap to meter).

また、料金については、入力トークンが $0.042 / MTok(100万トークン)、出力トークンは課金対象外とされています。

Parallel Sampling

Parallel. Generates all outputs in a single query.

一般的なLLMは、前のTokenをもとに次のTokenを生成する逐次的(Sequential)なToken生成を行います。

一方、Jevでは、文字列をToken単位で順番に生成するのではなく、複数のDecisionを並列(Parallel)に生成します。

# LLM 

Token 
↓ 
Token
↓
Token
↓
Token
# Jev

State 
↓ 
┌───────────┬───────────┬───────────┐ 
Decision A   Decision B  Decision C 
└───────────┴───────────┴───────────┘ 

              Parallel

TypeSafe AIでは、このParallel Samplingを、Jevが高速かつ効率的にDecisionを生成できる理由の1つとして挙げています。

また、Jevは各Decisionについて確率を並列に出力するため、複数の判断を含むWorkflowでも利用しやすい設計になっています。

Noul / Choice / Score

TypeSafe AIでは、Jevを利用したWorkflowで、主に以下の3種類のDecisionを使用しています。

We use three types: Noul, yes or no; Choice, one option among several; Score, a level on a scale.

種類 概要 出力
Noul Yes / Noの判断 YesとなるProbability
Choice 複数候補から1つを選択 各候補のProbabilityとConfidence
Score 尺度上で評価 Score、各LevelのProbability、Confidence

Noul

Noulは、Yes / Noで答えられる判断に利用します。
例えば、以下のような質問です。

この問い合わせは緊急対応が必要か?

Jevは単純にtrueやfalseを返すのではなく、YesであるProbabilityを返します。

Choice

Choiceは、あらかじめ定義した複数の候補から1つを選択する判断に利用します。

例えば、問い合わせを以下のカテゴリに分類できます。

  • billing
  • technical
  • refund
  • other

Jevは選択結果に加えて、各候補のProbabilityやConfidenceを返します。

Score

Scoreは、順序を持つ尺度上で評価する判断に利用します。

例えば、問い合わせの重要度を以下のような尺度で評価できます。

  • Low
  • Medium
  • High
  • Critical

JevはScoreだけでなく、各LevelのProbabilityとConfidenceを返します。

ユースケース

AI-Powered Workflows / smart if-statements.

TypeSafe AIでは、Jevの代表的な用途として、ソフトウェア内のAI-Powered Workflowやsmart if-statementsを挙げています。

具体的には、以下のような処理です。

  • classify:分類する
  • route:処理先を選択する
  • score:評価する
  • extract:情報を抽出する
  • branch:処理を分岐する

通常のif文では条件を明確なルールとして記述する必要があります。

一方、自然言語や曖昧な状態を扱う場合、条件をすべてルール化することが難しいケースがあります。

Jevでは、このような判断部分をDecisionとして実行し、その結果を通常のプログラムから利用できます。

Application State
        ↓
       Jev
        ↓
    Decision
        ↓
 ┌──────┴──────┐
処理A          処理B

また、TypeSafe AIでは以下のような用途も挙げています。

  • 大量データに対する分類・評価
  • リアルタイムアプリケーション
  • LLMの出力やReasoningの評価
  • Guardrail
  • Jailbreakの検出

実践

ここからは、実際にJevを利用して動作を確認します。

今回は、以下の流れで試します。

  1. TypeSafe AIのアカウントを作成する
  2. API Keyを作成する
  3. Noul / Choice / Scoreを試す
  4. 複数のDecisionを一度に実行する
  5. Confidenceを利用して処理を分岐する

1. アカウントを作成する

Jevは2026年9月15日の公開時点ではEarly Accessとして提供が開始されています。
Waitlist に登録してからアカウント作成になります

  1. しばらくすると、「TypeSafe AI: Your account is ready」の件名で、下記のようなメールが届きます

  2. 次に進めていくと、「Welcome to TypeSafe — confirm your email」が届きます

  3. 「Confirm email」をクリックします

  4. ブラウザが開くので、「Continue to TypeSafe」をクリックします

  5. 次に、Emailを入力し「Continue」をクリックします

  6. 「Sign in to TypeSafe」の件目で、下記のようなメールが届きます。「Sign in」をクリックします

  7. ブラウザが開くので、「Continue to TypeSafe」をクリックします

  8. Review our termsが開くので、内容を理解したうえで、「I have read and agreed to the Master Customer Agreement, Terms of Use, and Data Processing Addendum.」にチェックをし、「Continue」をクリックします

  9. 下記を入力して「Next」をクリックします

    • Name
    • Email
    • What's your primary area of work?
  10. 「Enter Console」をクリックします

  11. ポップアップが表示されます。「Yes」をクリックします

  12. Consoleが表示されました

2. API Keyを作成する

次に、JevのAPIを利用するためのAPI Keyを作成します。

  1. ナビゲーションペインの「API Keys」をクリックします

  2. 「Create API Key」をクリックします

  3. API Keyの名前を入力し、「Create key」をクリックします

  4. API Keyが作成されるので安全なところで保存してください

  5. 作成後はAPI Keysで確認できます
    image.png

API KEYを環境変数に設定する

  1. 作成したAPI Keyを環境変数 TYPESAFE_API_KEY に設定します

    export TYPESAFE_API_KEY="<作成したAPI Key>"
    
  2. 以降のAPI実行では、この環境変数を利用します

3. Jevを実行する

API Keyの準備ができたので、Jevを実行します。
まずは、最小構成のサンプルを使って、APIへ正常にアクセスできることを確認します。

SDKをインストールする

  1. 以下のコマンドで、TypeSafe AIのPython SDKをインストールします
pip install typesafe-sdk

サンプルコードを作成する

公式SDKのQuickstartをベースにしています。

今回は、以下の問い合わせを分類してみます。

I was charged twice. Please fix this ASAP.
  1. sample.pyを作成します
  2. 以下を記載します
from typesafe_sdk import Choice, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        state={
            "document": "I was charged twice. Please fix this ASAP."
        },
        questions={
            "category": Choice(
                instructions="What is this ticket about?",
                criteria={
                    "billing": None,
                    "technical": None,
                    "other": None,
                },
            ),
        },
    )

print(response.choices["category"].choice)

実行する

  1. 以下のコマンドで実行します
    % python sample.py                                                                  
    billing
    % 
    

入力した問い合わせは二重請求に関する内容だったため、billing に分類されました。
今回のコードでは、以下の3つの候補をあらかじめ定義しています。

  • billing
  • technical
  • other

Jevは自由な文章を生成するのではなく、定義された候補の中からDecisionを返します。

State
  ↓
"I was charged twice. Please fix this ASAP."
  ↓
Jev
  ↓
Choice
  ↓
billing

4. Noul / Choice / Scoreを試す

次に、Jevで利用できる3種類のDecisionである、Noul / Choice / Scoreを試します。
今回は、同じ問い合わせに対して以下を判断します。

種類 確認する内容
Noul 緊急性があるか
Choice 問い合わせのカテゴリ
Score 顧客の不満度

サンプルコードを作成する

  1. decision.py を作成します
  2. 以下を記載します
    from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
    
    state = {
        "ticket": "I was charged twice and need the duplicate refunded today."
    }
    
    with TypeSafeClient() as client:
        response = client.system_one(
            state=state,
            questions={
                "is_urgent": Noul(
                    instructions="Does the ticket explicitly communicate time pressure?"
                ),
                "category": Choice(
                    instructions="What is the customer's main request?",
                    criteria={
                        "refund": "The customer wants money returned.",
                        "technical_help": "The customer needs a bug or integration fixed.",
                        "information": "The customer is asking for information only.",
                        "other": "None of the other options clearly fits.",
                    },
                ),
                "frustration": Score(
                    instructions="How frustrated does the customer appear?",
                    criteria=[
                        "Calm and neutral",
                        "Concerned but civil",
                        "Very angry or using strong language",
                    ],
                ),
            },
        )
    
    print("Noul:")
    print(response.nouls["is_urgent"].noul)
    
    print("\nChoice:")
    print(response.choices["category"].choice)
    print(response.choices["category"].probabilities)
    print(response.choices["category"].confidence)
    
    print("\nScore:")
    print(response.scores["frustration"].score)
    print(response.scores["frustration"].probabilities)
    print(response.scores["frustration"].confidence)
    

実行する

  1. 以下のコマンドで実行します

    % python decision.py
    Noul:
    0.96
    
    Choice:
    refund
    {'technical_help': 0.0, 'refund': 1.0, 'other': 0.0, 'information': 0.0}
    1.0
    
    Score:
    0.96
    {0: 0.04, 1: 0.96, 2: 0.0}
    0.94
    % 
    
    • Noul:0.96
      • Noulは、Yes / Noで答えられるQuestionを定義する型
      • is_urgentについて 96%の確率でYesと判断しています
      • 「時間的な切迫感が明示されている」がYesであるProbability = 96% という意味
    • Choice:
          refund
          {'technical_help': 0.0, 'refund': 1.0, 'other': 0.0, 'information': 0.0}
          1.0
      
      • Choice は、あらかじめ用意した複数候補の中から、どれが最も当てはまるかをJevに選ばせるDecisionです
      • 今回のStateは、「I was charged twice and need the duplicate refunded today.」
      • に対して、「What is the customer's main request?」という質問
      • それぞれのカテゴリーに最もハマるかをJevが判断している
      • refundが1.0と最も高い結果になりました
    • Score:
          0.96
          {0: 0.04, 1: 0.96, 2: 0.0}
          0.94
      
      • Score は、複数候補から1つを選ぶのではなく、順序のある尺度のどの位置に近いかを評価するDecision
      • 今回のStateは、「I was charged twice and need the duplicate refunded today.」
      • に対して、「How frustrated does the customer appear?」という質問
      • 今回の尺度では下記のようになっている
        0 = Calm and neutral
        1 = Concerned but civil
        2 = Very angry
        
      • Concerned but civil(0.96)が選ばれました

5. Confidenceを利用して処理を分岐する

最後に、Jevが返すConfidenceを利用して、処理を分岐してみます。

今回は、問い合わせカテゴリの判定結果について、

  • Confidenceが0.9以上の場合は自動処理
  • Confidenceが0.9未満の場合はHuman Review

とします。

サンプルコードを作成する

  1. confidence.py を作成します
  2. 以下を記載します
from typesafe_sdk import Choice, TypeSafeClient

state = {
    "ticket": "I was charged twice and need the duplicate refunded today."
}

with TypeSafeClient() as client:
    response = client.system_one(
        state=state,
        questions={
            "category": Choice(
                instructions="What is the customer's main request?",
                criteria={
                    "refund": "The customer wants money returned.",
                    "technical_help": "The customer needs a bug or integration fixed.",
                    "information": "The customer is asking for information only.",
                    "other": "None of the other options clearly fits.",
                },
            ),
        },
    )

result = response.choices["category"]

print("Decision:")
print(result.choice)

print("\nProbability:")
print(result.probabilities)

print("\nConfidence:")
print(result.confidence)

if result.confidence >= 0.9:
    print("\nAction: Auto Process")
else:
    print("\nAction: Human Review")

実行する

  1. 以下のコマンドで実行します
    % python confidence.py
    Decision:
    refund
    
    Probability:
    {'refund': 1.0, 'other': 0.0, 'technical_help': 0.0, 'information': 0.0}
    
    Confidence:
    1.0
    
    Action: Auto Process
    % 
    
    • 入力した問い合わせ「I was charged twice and need the duplicate refunded today.」
    • 結果は、「refund」が選択されています
    • Jevは今回の問い合わせについて、refund であるProbabilityを 1.0 としています
    • Actionでは、confidenceが 90%(0.9)以上であれば、Auto Processとしています。今回もこちらに該当しています

考察

Jevを実際に試してみて、一般的なLLMとは異なり、文章を生成するのではなく、アプリケーション内で利用するDecisionを返すことに特化したモデルであることを確認できました。

特に、Noul / Choice / Scoreを使い分けることで、アプリケーションが必要とする判断を小さなQuestionとして定義できる点が特徴的だと感じました。

  • Noul:Yes / Noで判断する
  • Choice:複数の候補から選択する
  • Score:順序のある尺度上で評価する

今回のハンズオンでは、ChoiceのConfidenceが 0.9 以上の場合は自動処理、それ未満の場合はHuman Reviewとする処理を実装しました。

AIに処理全体を任せるのではなく、曖昧な判断をJevに任せ、その結果を通常のアプリケーションコードで制御するという使い方が、Jevの特徴の1つだと考えます。

次回は、複数のDecisionを組み合わせてWorkflowを作ることを試してみてたいです。

参考

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?