はじめに
こんにちは、私は機械知能研究室に所属し、CV分野の研究をしています
機械知能研究室ではロボットの視覚機能や自律的に行動するための知能システムについての研究に取り組んでいます (ホームページはこちら)
本題:今回は2026年2月に登場したMoonshine Voice[1]について、AIの研究をしている学生が知っておくべきレベルでサクッと理解できるように感想と共にまとめてみました
Moonshine Voiceとは
一言でまとめるなら"音声を入力し、テキストにして出力するモデル"のツールキットです
時系列整理
- 2024年:リアルタイム音声認識のMoonshine[2]というモデルが提案
- 2026年:レイテンシが重要なアプリケーション向けのMoonshine v2[3]が提案
- 2026年(おそらく):Moonshine v2を使用した、Moonshine Voice[1]が公開
つまり、Moonshine voiceはMoonshine(v2をベースとした)モデルを使用したツールキット
ここで、以降の本記事では
- Moonshineについて
- アプリケーション向けのMoonshine v2について
- Moonshine Voiceの紹介
という流れで記載します
1 Moonshineの概要

アーキテクチャ(Moonshineのarxiv論文[2]のFigure 3から引用)
1.1 アーキテクチャ・流れ
- 音声を畳み込み処理で圧縮
- Transformer[4]のEncoderに入力しDecoderからtoken IDを出力
- token IDをBPE[5](トークナイザー)でテキストを生成
トークナイザー:token IDからテキスト(またはその逆)にするもの
1.2 モデルの特徴
Transformerには、一般的な位置埋め込みではなく、Rotary Position Embedding(RoPE)[6]を採用
RoPE:埋め込みベクトルに回転行列をかける位置埋め込み方法
1.3 性能と私の感想
WERという指標[8]において、同サイズ(tiny, base)のWhisper[7]というモデルを上回る性能であった
感想:WERの値の差が小数点以下であり、WERの値はどれだけ下がったら凄いものかが理解できていないので、どれほどすごいかはわかりませんでした
Moonshine v2の概要
Moonshine v2の論文[3]から説明
図がなかったので、テキストで流れを解説
基本的な構造はMoonshineと同様だと思われます
- オーディオプリプロセッサ(前処理)
- Transformer Encoder
- アダプタ(位置埋め込み)
- Transformer Decoder
- トークナイザー(トークンIDからテキスト生成)
特徴
Sliding-window attention encodersと、No positional embeddings(ergodc encoder)を採用
これにより、6倍のサイズのモデルと同等の品質で、大幅に高速に動作
Sliding-window attention encoders
通常のアテンション:全てのフレームを入力後、アテンションを計算しテキストを生成
改善:任意のフレーム数のみでアテンションを計算し、随時テキストを生成
Full-Attention Encoderとの比較(MoonshineV2のarxiv論文[3]のFigure 2から引用)
ergodic encoder
Encoderに位置埋め込みをしない
(入力の音声には順序があり、Encoderの前の畳み込み処理でも順序情報をもつ)
Encoderの出力をアダプタで位置埋め込みをし、Decoderに入力
1.3 性能と私の感想
Time-to-first-token(TTFT)と入力の長さの関係(MoonshineV2のarxiv論文[3]のFigure 5から引用)
TTFT:音声が入力されてから最初のテキストが出力されるまでの時間
Moonshine V2は入力の長さに関係なく同じTTFTである
(入力が長くても最初のテキスト生成までの時間が変わらない)
感想:スライディングウィンドウアテンションはとても魅力的に感じた
また、位置埋め込みをしていないEncoderの出力をアダプタで位置埋め込みしても良いのかと疑問に思った
Moonshine Voiceの紹介
Moonshine VoiceのGitHubが公開されている
https://github.com/moonshine-ai/moonshine?tab=readme-ov-file
英語版はMITライセンス、その他の言語は非営利ライセンス(Moonshine comunity Licence)の下で公開
iOS, MacOS, Linuxなど、様々なプラットフォームで利用可能
実際に私の環境(MacOS M1pro)で実行を試みましたが、OSが古く、アップデートが必要なので断念しました(MacOS 14以上で動くと思います)
README.mdを見る限りでは簡単に実行できそうです
終わりに
私もこの分野に関しては全くの初学者ですが、この記事を書くことを通してある程度理解できたと思います(読者の方も同じであれば幸いです)
ある程度理解したことにより、このツールキットやモデルに今後どのような発展があるか楽しみです
参考文献・参考にしたサイト
[1]Moonshine VoiceのGitHub(https://github.com/moonshine-ai/moonshine)
[2]Moonshineのarxiv論文(Moonshine: Speech Recognition for Live Transcription and Voice Commands)
@misc{jeffries2024moonshinespeechrecognitionlive,
title={Moonshine: Speech Recognition for Live Transcription and Voice Commands},
author={Nat Jeffries and Evan King and Manjunath Kudlur and Guy Nicholson and James Wang and Pete Warden},
year={2024},
eprint={2410.15608},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2410.15608},
}
[3]Moonshine V2のarxiv論文(Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications)
@misc{kudlur2026moonshinev2ergodicstreaming,
title={Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications},
author={Manjunath Kudlur and Evan King and James Wang and Pete Warden},
year={2026},
eprint={2602.12241},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.12241},
}
[4]Transformerのarxiv論文(Attention Is All You Need)
@misc{vaswani2023attentionneed,
title={Attention Is All You Need},
author={Ashish Vaswani and Noam Shazeer and Niki Parmar and Jakob Uszkoreit and Llion Jones and Aidan N. Gomez and Lukasz Kaiser and Illia Polosukhin},
year={2023},
eprint={1706.03762},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/1706.03762},
}
[5]BPEに関する論文(Neural Machine Translation of Rare Words with Subword Units)
@misc{sennrich2016neuralmachinetranslationrare,
title={Neural Machine Translation of Rare Words with Subword Units},
author={Rico Sennrich and Barry Haddow and Alexandra Birch},
year={2016},
eprint={1508.07909},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/1508.07909},
}
[6]RoPEを提案している論文(RoFormer: Enhanced Transformer with Rotary Position Embedding)
@misc{su2023roformerenhancedtransformerrotary,
title={RoFormer: Enhanced Transformer with Rotary Position Embedding},
author={Jianlin Su and Yu Lu and Shengfeng Pan and Ahmed Murtadha and Bo Wen and Yunfeng Liu},
year={2023},
eprint={2104.09864},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2104.09864},
}
[7]Wisperのarxiv論文Robust Speech Recognition via Large-Scale Weak Supervision
@misc{radford2022robustspeechrecognitionlargescale,
title={Robust Speech Recognition via Large-Scale Weak Supervision},
author={Alec Radford and Jong Wook Kim and Tao Xu and Greg Brockman and Christine McLeavey and Ilya Sutskever},
year={2022},
eprint={2212.04356},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2212.04356},
}
[8]【0から学ぶAI】第317回:音声認識モデルの評価指標 〜Word Error Rate(WER)などの指標を説明
Whisperを超える精度のリアルタイム文字起こしローカルAI「Moonshine Voice」、日本語にも対応(生成AIクローズアップ)

