中国に出張中でもノートパソコンで利用できるAIが欲しい!
文章の体裁を整えたりするためにCopilotを使いたいが、中国などでは使えないらしい(サービス提供外とか出るそうな)。
打ち合わせ内容のまとめなど、文章の体裁を整えてもらうのにAI機能を使いたいけれど、通常の個人向けのAIサービスを使われると情報漏洩になってしまうので、ノートパソコンで実行可能なローカルLLMを準備することに。
ノートパソコン環境
CPU:Intel 11th Gen Intel(R) Core(TM) i5-1145G7
MEM:16.0 GB
OS:Windows11 Enterprise 24H2
Intelの統合GPUなのでOpenVINOを利用したいところだけど、一般的なローカルLLMの実行アプリはOpenVINOには対応していないらしい。
GeminiによるとPythonで作ればできるよ!ということだったので、簡単なGUIを作ってpyinstallerで配布することに。
導入したPythonパッケージ
pip install openvino-genai optimum-intel datasets huggingface-hub==1.33.0RC1 pyinstaller
| パッケージ | インストールされたバージョン |
|---|---|
| openvino-genai | 2026.4.1.0 2636 |
| optimum | 2.3.0 |
| optimum-intel | 2.2.0 |
| datasets | 5.1.0 |
| huggingface_hub | 1.33.0rc1 |
| pyinstaller | 6.22.3 |
huggingface_hubを最新にするとoptimumでOPENVINOへの変換ができなかったので、これだけバージョンを指定して導入しています。
ローカルLLMモデルの検討
日本語の文章を扱うことが中心の予定なので、日本語性能とノートパソコンでも利用できるレスポンスを重視。
メモリが16GBなので、他のアプリの利用も考えると8Bクラスがギリギリかも?というのはGeminiの回答。
| Hugging Face model | 備考 |
|---|---|
Qwen/Qwen2.5-7B-Instruct |
OpenVINOに変換。快適に動作したが日本語が不自由。 |
tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2 |
東京科学大学で日本語の追加学習されたQwen3。日本語が優れている。OpenVINOに変換必要。やや重め。 |
OpenVINO/LFM2.5-8B-A1B-int4-ov |
A1Bでレスポンスが良い。性能は高いが日本語は問題ないが味気ない。 |
OpenVINO/gemma-4-E4B-it-int4-ov |
そこそこのレスポンスで十分に流暢な日本語。今回は不要だがマルチモーダル対応 |
OpenVINO公式にて変換済みの中から日本語が流暢だったGemma-4-E4Bを選択した。
モデルのダウンロードと変換
- モデルダウンロードとOpenVINO変換
Qwen2.5-7B-InstructをOpenVINO変換した際のコマンド
optimum-cli export openvino --model Qwen/Qwen2.5-7B-Instruct --weight-format int4 ./qwen25_7b_openvino_int4 - モデルのダウンロード
Hugging FaceのOpenVINO公式からOpenVINO変換済みのGemma-4-E4Bをダウンロードしたコマンド
hf download OpenVINO/gemma-4-E4B-it-int4-ov
コード
ほぼほぼGeminiなどに作ってもらったが、試行錯誤の過程でいくつか躓いた点を記載します。
-
pipe = ov_genai.VLMPipeline(str(MODEL_DIR), "GPU", config=config)
最初にテストしていたQwen2.5-7B-Instructではov_genai.LLMPipelineで動作したが、Gemma-4はマルチモーダルのためか、VLMPipelineを利用する模様。
GeminiなどにはCPUを推されるが、GPUとした方が推論が早かったので、GPUをおすすめします。 - ディスプレイドライバーが古い問題
これもQwen2.5-7B-Instructでは問題なかったが、Gemma-4ではGPUではエラーとなって動作せず。標準のディスプレイドライバーが古くOpenVINO対応にバグがあるようで、ドライバーバージョン 32.0.101.7092にしたところ、GPUで正しく推論できるようになった。
localllm_gui.py
import openvino_genai as ov_genai
import tkinter as tk
from tkinter import scrolledtext
import threading
from pathlib import Path
import sys
if getattr(sys, 'frozen', False):
BASE_DIR = Path(sys.executable).parent
else:
BASE_DIR = Path(__file__).parent
MODEL_DIR = BASE_DIR / "gemma-4-E4B-it-int4-ov"
# コンパイルの効率化(キャッシュ)を設定します
config = {
"CACHE_DIR": "./openvino_cache"
}
try:
# VLMPipelineの初期化(GPUを指定)
pipe = ov_genai.VLMPipeline(str(MODEL_DIR), "GPU", config=config)
print("GPUでの初期化に成功しました。")
except Exception as e:
print(f"GPUでの初期化に失敗しました。CPUに切り替えます: {e}")
pipe = ov_genai.VLMPipeline(str(MODEL_DIR), "CPU")
root = tk.Tk()
root.title("Local LLM GUI (Gemma-4)")
# --- プロンプト入力部分 ---
prompt_label = tk.Label(root, text="プロンプトを入力してください:")
prompt_label.pack(anchor="w", padx=5, pady=5)
input_frame = tk.Frame(root)
input_frame.pack(fill="x", padx=5, pady=5)
prompt_text = tk.Text(input_frame, height=5, width=70, wrap=tk.WORD)
prompt_text.pack(side="left", fill="both", expand=True)
prompt_scrollbar = tk.Scrollbar(input_frame, orient="vertical", command=prompt_text.yview)
prompt_scrollbar.pack(side="left", fill="y")
prompt_text.config(yscrollcommand=prompt_scrollbar.set)
prompt_text.focus()
generate_button = tk.Button(input_frame, text="生成開始")
generate_button.pack(side="left", padx=5, pady=0, anchor="n")
# --- 出力表示部分 ---
output_frame = tk.Frame(root)
output_frame.pack(fill="both", expand=True, padx=5, pady=5)
output_text = scrolledtext.ScrolledText(output_frame, height=20, width=80, wrap=tk.WORD)
output_text.pack(side="left", fill="both", expand=True)
def copy_output():
text = output_text.get("1.0", tk.END).strip()
if text:
root.clipboard_clear()
root.clipboard_append(text)
copy_button = tk.Button(output_frame, text="コピー", command=copy_output)
copy_button.pack(side="left", padx=5, pady=5, anchor="n")
# --- 生成処理 ---
system_prompt = ""
def generate_response():
prompt = prompt_text.get("1.0", tk.END).strip()
if not prompt:
return
full_prompt = system_prompt + prompt
output_text.delete("1.0", tk.END)
first_token_received = False
output_text.insert(tk.END, "考え中...\n")
def streamer(subword):
nonlocal first_token_received
def update_text():
nonlocal first_token_received
if not first_token_received:
current_text = output_text.get("1.0", tk.END)
if "考え中..." in current_text:
output_text.delete("1.0", "2.0")
first_token_received = True
output_text.insert(tk.END, subword)
output_text.see(tk.END)
root.after(0, update_text)
def run_generation():
try:
pipe.generate(full_prompt, max_new_tokens=2048, streamer=streamer)
root.after(0, lambda: output_text.insert(tk.END, "\n"))
except Exception as err:
# lambdaのスコープバグを回避するため、引数 msg に初期値として代入
error_msg = f"\n[Error] {err}\n"
root.after(0, lambda msg=error_msg: output_text.insert(tk.END, msg))
threading.Thread(target=run_generation, daemon=True).start()
generate_button.config(command=generate_response)
root.mainloop()
pyinstallerでEXE化
モデルフォルダはマニュアルでコピペする。
pyinstaller --icon localllmgui.ico --onedir --windowed --collect-all openvino --collect-all openvino_genai --collect-all openvino_tokenizers localllm_gui.py
感想など
- Qwen2.5-7B-Instructは、OpenVINOでの動作テストとしては良好だったが、日本語が不自由で常用できる感じではなかった。
- Gemma-4-E4Bはエッヂ向けの小さなモデルなのに流暢な日本語で破綻もなく、安心して利用できる。今回は利用していないがマルチモーダル対応なので他にも便利に使えそう。
- 統合GPU + OpenVINOでローカルLLMはどこまで戦えるのか?というのが個人的には注目点だったが、今回の日本語文章の手直し程度であればストレスも少なく十分に利用できそう。偶然ではあるがIntel 11世代はOpenVINOでのAI推論に向いていたようでローカルでも結構戦える印象。
- Pythonでサンプルも多いのか無課金のGeminiで大きな問題なく生成できた。エラーについてもエラーコード丸投げで修正案を提示してもらえ楽ちん。
- C#などで作成しWindowsネイティブとした方がメモリやアプリ起動が早く、さらにストレスがないだろうとは思ったが、本業ではなく
お金ももらえないのでpyinstallerでお茶を濁した。
おまけ
勢いで文字起こしツールもGeminiに作ってもらいました。
transcribe.py
import os
import threading
import time
import tkinter as tk
from tkinter import filedialog, messagebox
from tkinter import ttk
import librosa
from huggingface_hub import snapshot_download
import openvino_genai as ov_genai
class TranscribeApp:
def __init__(self, root):
self.root = root
self.root.title("OpenVINO 音声文字起こしツール")
self.root.geometry("600x320")
self.root.resizable(False, False)
# 変数管理
self.audio_file_path = tk.StringVar()
self.status_text = tk.StringVar(value="音声ファイルを選択してください。")
self.progress_value = tk.DoubleVar(value=0.0)
# モデル設定
self.repo_id = "OpenVINO/whisper-large-v3-turbo-int8-ov"
self.model_dir = "whisper-large-v3-turbo-int8-ov"
self.device = "GPU" # 第11世代CPUのIris Xe用にGPUを指定
self.create_widgets()
def create_widgets(self):
# 1. ファイル選択エリア
file_frame = tk.LabelFrame(self.root, text=" 1. 音声ファイルの選択 ", padx=10, pady=10)
file_frame.pack(fill="x", padx=15, pady=10)
entry_path = tk.Entry(file_frame, textvariable=self.audio_file_path, width=55)
entry_path.pack(side="left", padx=(0, 10), expand=True, fill="x")
btn_browse = tk.Button(file_frame, text="参照...", command=self.browse_file, width=10)
btn_browse.pack(side="right")
# 2. 実行エリア
action_frame = tk.Frame(self.root, pady=10)
action_frame.pack(fill="x", padx=15)
self.btn_start = tk.Button(
action_frame,
text="変換開始",
command=self.start_transcription_thread,
bg="#0078D4",
fg="white",
font=("Arial", 11, "bold"),
height=2
)
self.btn_start.pack(fill="x")
# 3. ステータス・プログレス表示エリア
status_frame = tk.LabelFrame(self.root, text=" 2. 処理ステータス ", padx=10, pady=10)
status_frame.pack(fill="both", expand=True, padx=15, pady=(5, 15))
lbl_status = tk.Label(status_frame, textvariable=self.status_text, wraplength=550, anchor="w", justify="left")
lbl_status.pack(fill="x", pady=(0, 10))
self.progress_bar = ttk.Progressbar(
status_frame,
orient="horizontal",
mode="determinate",
variable=self.progress_value
)
self.progress_bar.pack(fill="x")
def browse_file(self):
file_path = filedialog.askopenfilename(
title="音声ファイルを選択",
filetypes=[("Audio Files", "*.m4a *.mp3 *.wav"), ("All Files", "*.*")]
)
if file_path:
self.audio_file_path.set(file_path)
self.status_text.set("変換ボタンを押すと文字起こしを開始します。")
self.progress_value.set(0.0)
def update_progress(self, value):
self.root.after(0, lambda: self.progress_value.set(value))
def update_status(self, text):
self.root.after(0, lambda: self.status_text.set(text))
def start_transcription_thread(self):
path = self.audio_file_path.get()
if not path or not os.path.exists(path):
messagebox.showerror("エラー", "有効な音声ファイルを選択してください。")
return
self.btn_start.config(state="disabled")
threading.Thread(target=self.process_transcription, args=(path,), daemon=True).start()
def process_transcription(self, audio_path):
try:
# 1. モデルのダウンロードチェック
if not os.path.exists(self.model_dir):
self.update_status("[*] 初回起動: AIモデルをダウンロード中 (数分かかります)...")
snapshot_download(repo_id=self.repo_id, local_dir=self.model_dir)
# 2. 音声ファイルのロードと長さ計測
self.update_status("[*] 音声ファイルを読み込んでいます...")
raw_speech, sr = librosa.load(audio_path, sr=16000, mono=True)
# 3. パイプラインの初期化
self.update_status(f"[*] AIエンジンを初期化中 (デバイス: {self.device})...")
pipe = ov_genai.WhisperPipeline(self.model_dir, self.device)
config = pipe.get_generation_config()
config.language = "<|ja|>"
config.task = "transcribe"
config.return_timestamps = True
# 4. 音声を30秒ごとに分割してループ処理
chunk_length_secs = 30
chunk_length_samples = chunk_length_secs * sr
total_samples = len(raw_speech)
# 出力用のテキストを保持するリスト
all_chunks_text = []
self.update_status("[*] 文字起こしを実行中... (Intel GPU加速)")
start_time = time.time()
# 30秒ごとにシーク(切り出し)しながら処理
for start_sample in range(0, total_samples, chunk_length_samples):
# 30秒分の音声データを切り出す
end_sample = min(start_sample + chunk_length_samples, total_samples)
audio_chunk = raw_speech[start_sample:end_sample].tolist()
# 現在の音声内の再生位置(秒)
current_sec_offset = start_sample / sr
# 30秒のデータを推論
result = pipe.generate(audio_chunk, config)
# 取得したテキストの時間情報を全体の時間にオフセットする
for chunk in result.chunks:
actual_start = current_sec_offset + chunk.start_ts
actual_end = current_sec_offset + chunk.end_ts
start_min, start_sec = int(actual_start // 60), int(actual_start % 60)
end_min, end_sec = int(actual_end // 60), int(actual_end % 60)
timestamp_str = f"[{start_min:02d}:{start_sec:02d} -> {end_min:02d}:{end_sec:02d}]"
all_chunks_text.append(f"{timestamp_str} {chunk.text}")
# 進捗率を更新 (0.0 〜 100.0)
progress = (end_sample / total_samples) * 100
self.update_progress(progress)
end_time = time.time()
# 5. 同一フォルダへのテキスト出力保存
base_path, _ = os.path.splitext(audio_path)
txt_output_path = base_path + ".txt"
with open(txt_output_path, "w", encoding="utf-8") as f:
for line in all_chunks_text:
f.write(line + "\n")
# 完了メッセージ
elapsed = end_time - start_time
success_msg = f"【完了】保存先:\n{os.path.basename(txt_output_path)}\n(処理時間: {elapsed:.1f}秒)"
self.update_status(success_msg)
self.root.after(0, lambda: messagebox.showinfo("完了", "文字起こしが正常に完了しました!"))
except Exception as e:
error_message = str(e)
self.update_status(f"【エラー発生】 {error_message}")
self.root.after(0, lambda msg=error_message: messagebox.showerror("エラー", f"処理中にエラーが発生しました:\n{msg}"))
finally:
self.root.after(0, lambda: self.btn_start.config(state="normal"))
if __name__ == "__main__":
root = tk.Tk()
app = TranscribeApp(root)
root.mainloop()