0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

ローカル画像生成で「Corporate Memphis」を本物っぽく出す。Qwen-Imageのプロンプト/ネガティブ/img2img実装知見

0
Last updated at Posted at 2026-08-01

ローカル(Qwen-Image)で生成した高品質なCorporate Memphis

ブログのヒーロー画像やOGP画像を、クラウドの画像生成API(Gemini等)からローカルモデルに置き換えられるか。この検証を先日おこないました(全体像は本家記事にまとめています)。

本家記事: ローカル画像生成AIはGeminiの代わりになるか。FLUX・Qwen-Image・Playground v2.5を商用利用前提で実機比較した

その中で一番てこずったのが「Corporate Memphis(コーポレート・メンフィス)」スタイルでした。フラットで手足が長い、あのビッグテック系イラストです。ローカルモデルだと最初はまったく別物になり、作り込んでもクオリティが上がらない。

この記事では、そのCorporate Memphisをローカル(Qwen-Image)でGemini同等以上の品質に持っていくまでの実装知見を、プロンプト・ネガティブプロンプト・img2img・ハーネスのコードまで含めて共有します。

検証環境

  • GPU: Radeon RX 7900 XTX 24GB(ROCm)
  • モデル: Qwen-Image(Q6_K GGUF量子化・Apache-2.0で商用OK)、diffusers
  • 生成解像度: 1216×640、seed固定で比較

※速度・VRAMはこの構成での実測値です。環境で大きく変わります。

つまずき① 「スタイル名を指定しても別物になる」

最初の失敗は、同じ「Corporate Memphis」を指定しているのに、Geminiの既存ヒーローと構図からして別物になったこと。原因はモデルの実力ではなくプロンプトの構図テンプレでした。

  • Geminiヒーロー: 「大きなメインオブジェクト中心+周囲に小さな人物を配置」
  • こちらのテスト: 「複数人の人物が主役で作業している場面」

構図テンプレを揃えるだけでGeminiにグッと近づきます。

構図テンプレ・プロンプト(日本語訳)

ビッグテック系のフラットなベクターイラスト、
中央に大きなノートPC(シンプルで見やすいダッシュボードを表示)、
その周りに全身の小さな人物を数人、リラックスした自然な姿勢で配置、
...
濃紺の背景、横長の構図、画像内に文字は一切入れない

構図テンプレ・プロンプト(原文 - 英語)

big-tech corporate flat vector illustration,
a large laptop computer displaying a simple clean app dashboard in the center,
several small full-body human figures scattered around the laptop in relaxed graceful poses,
...
dark navy blue background, wide horizontal composition, absolutely no text anywhere

ポイント: several small ... figures ... around(メイン中心+周囲に小さな人物)が、GeminiのCMと構図を揃える肝です。

Corporate Memphisの品質改善: Gemini目標 / 旧プロンプト(別物) / 構図テンプレ(地味) / 服・髪・参照画像で追い込んだ高品質版

つまずき② 「単色の棒人間になる」

構図を揃えても、今度は人物が単色ボディの棒人間になり、地味でした。本物のCorporate Memphisは、顔は最小限(点目 or 無表情)でも服(上下で色違い)・髪・多様な肌色がちゃんとあります。ここを明示していなかったのが原因です。

ポジティブ側に「服・髪・多様性」を足します。

服・髪・多様性のプロンプト(日本語訳)

各人物はフラットな単色の髪を持つ、
上(シャツ)と下(ズボン)で色の違うシンプルな服を着ている、
多様な肌の色、
小さくシンプルな丸い頭・無表情で細かい造作は無し、
上品でゆるやかに曲がった長い手足、
輪郭線・陰影・テクスチャの無い、なめらかでベタ塗りのフラットな塗り

服・髪・多様性のプロンプト(原文 - 英語)

each person has simple flat solid-color hair
and wears simple flat clothing with a distinct shirt color and separate trouser color,
diverse skin tones,
tiny simple round heads with blank minimal faces and no detailed features,
elegant gently curved long limbs,
smooth clean flat solid color fills with no outlines no shading and no texture

この"服・髪・多様な肌色"が無いと、単色の棒人間になります。

つまずき③ 「太い輪郭・大きい頭・写実の顔に逸脱する」

作り込むと、今度は太い黒アウトライン・大きな頭・表情つきの写実的な顔という別スタイル(ボールドなフラット挿絵)に流れます。ここはネガティブプロンプトが効きます。

ネガティブプロンプト(日本語訳)

文字・単語・アルファベット・タイポグラフィ・キャプション、
棒人間・単色ボディ・裸・服なし・ハゲ・髪なし、
縞・ストライプ・テクスチャ・ハッチング・肌の上の線・粒子ノイズ、
太い黒枠・大きすぎる頭、
書き込まれた顔・表情・細かい目・口・写実的な顔、
写実的・3Dレンダー・粘土・グラデーション・落ち影・クローズアップ・ポートレート

ネガティブプロンプト(原文 - 英語)

text, words, letters, typography, captions,
stick figure, single-color body, naked, no clothes, bald, no hair,
stripes, striped, texture, hatching, lines on skin, grain, noise,
thick black outlines, bold outlines, large heads, big heads,
detailed faces, facial expressions, detailed eyes, mouth, realistic face,
photorealistic, 3d render, clay, gradient, drop shadow, close-up, portrait

text/words/letters=文字化、stick figure/naked/bald/single-color body=棒人間化、stripes/texture=縞ノイズ、thick black outlines/large heads/detailed faces=太枠・大頭・書き込み顔、photorealistic/3d/clay/gradient/shadow=写実/立体化 を、それぞれ抑えています。特に stick figure / naked / bald / single-color body が棒人間化の抑制に効きます。

つまずき④ 「画像の中に "Memphis" という文字が出る」

地味なハマりどころ。プロンプトに "Corporate Memphis" と書くと、Qwenがそれを画像内の文字として描画してしまいます("Memphis" や "Alegia" などのロゴ文字が出る)。

対策:

  • スタイル名の固有名詞をプロンプトに書かない(big-tech corporate flat vector illustration 等で言い換える)
  • absolutely no text anywhere を付け、ネガに text, words, letters を入れる

仕上げ: 参照画像 img2img でスタイルを固定する

プロンプト単体でもかなり良くなりますが、きれいなCM参照画像から img2img(strength≈0.5)すると、さらに安定します。

  • 参照の構図を継承しつつ、人物を服・髪つきの高品質なものに置換できる
  • スタイル名由来の文字化も消える(テキストの無い参照に縛られるため)

diffusersでの実装(GGUF版Qwenのimg2imgパイプライン):

import torch
from PIL import Image
from diffusers import QwenImageImg2ImgPipeline, QwenImageTransformer2DModel, GGUFQuantizationConfig

transformer = QwenImageTransformer2DModel.from_single_file(
    "https://huggingface.co/city96/Qwen-Image-gguf/blob/main/qwen-image-Q6_K.gguf",
    quantization_config=GGUFQuantizationConfig(compute_dtype=torch.bfloat16),
    config="Qwen/Qwen-Image", subfolder="transformer", torch_dtype=torch.bfloat16,
)
pipe = QwenImageImg2ImgPipeline.from_pretrained(
    "Qwen/Qwen-Image", transformer=transformer, torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()

init = Image.open("ref_cm.png").convert("RGB").resize((1216, 640))
img = pipe(
    prompt=PROMPT, negative_prompt=NEG, image=init,
    strength=0.5, num_inference_steps=30, true_cfg_scale=4.0,
    height=640, width=1216,
    generator=torch.Generator("cpu").manual_seed(42),
).images[0]
img.save("out.png")

strengthは体感で 0.5前後が最適でした。0.65まで上げると躍動感は出ますが、スタイル名由来の文字が再発しやすくなります。

参照画像img2img(strength≈0.5)で仕上げた高品質なCorporate Memphis

まとめ(効いた順)

  1. 構図テンプレを揃える(メイン中心+周囲に小人物・長い手足・点目・no text)
  2. 服・髪・多様な肌色を明示(棒人間化の解消)
  3. ネガティブで逸脱を抑制(太輪郭・大頭・顔・文字)
  4. スタイル名をプロンプトに書かない(文字化回避)
  5. 参照画像から img2img(strength≈0.5) でスタイル固定(最も安定)

モデルの「実力不足」に見えていた品質差の大半は、プロンプト設計と参照画像の使い方で埋められました。ローカル(Apache-2.0のQwen-Image)でも、商用のヒーロー画像に十分使える品質のCorporate Memphisが出せます。

元記事

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?