ローカルLLM・小規模GPU環境における推論サービング設計とモデル選定を強みとしています。vLLM/SGLang/Ollama、量子化、KV Cache、並列推論、TTFT/ITLを実測し、用途・VRAM・レイテンシ・スループットから最適な実用構成を判断できます。
Location
TOKYO
Following Organizations
No Organizations you are following
Contributions
article is Stocked
article is Liked
article is Stocked
article is Stocked
posted an article
posted an article
article is Liked
posted an article
posted an article
article is Liked