0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

【第4回】商品ID1から派生URLを生成し、追加タグ情報を取得する

0
Posted at

【第4回】商品ID1から派生URLを生成し、追加タグ情報を取得する

はじめに

第3回で、一次ページから ProductRecord を作れるようになりました。
第4回では、その product_id1 を使って派生URLを作り、追加情報(specific_tag1)を取得してレコードへ統合します。

この回でできるようになること:

  • DETAIL_BASE_URL + product_id1 で派生URLを生成
  • 派生先ページから特定タグを安全に取得
  • 一次抽出データと派生抽出データを1つのレコードに統合

1. ProductRecord を拡張する(src/models.py)

まず、派生処理の結果を保持できるようにモデルを拡張します。

from dataclasses import dataclass
from typing import Optional


@dataclass
class ProductRecord:
    # 一次ページから取得
    product_id1: str
    product_id2: Optional[str]
    product_name: str
    issue1: Optional[str]
    issue2: Optional[str]
    source_url: str

    # 第4回で追加: 派生ページ関連
    derived_url: Optional[str] = None
    specific_tag1: Optional[str] = None

デフォルト値を None にしておくことで、段階実装中も壊れにくくなります。


2. 派生URLを作る関数を追加(src/extractor.py)

派生URL生成ロジックを関数化しておくと、あとで仕様変更に対応しやすいです。

from typing import Optional


def build_derived_url(detail_base_url: str, product_id1: str) -> Optional[str]:
    """
    DETAIL_BASE_URL と product_id1 を連結して派生URLを作る。

    product_id1 が空の場合はURLを作れないため None を返す。
    """
    if not product_id1:
        return None

    # 末尾スラッシュやクエリ付きURLにする場合は、業務仕様に合わせてここを調整。
    return f"{detail_base_url}{product_id1}"

3. 派生ページから追加タグを取得する関数を作る(src/extractor.py)

第3回の safe_text_content() を再利用して、取り方を統一します。

from playwright.sync_api import Page


def fetch_specific_tag1_from_derived_page(
    page: Page,
    derived_url: str,
    sel_derived_tag1: str,
) -> Optional[str]:
    """
    派生ページにアクセスし、specific_tag1 を取得する。

    取得できない場合は None を返し、上位処理で継続できる設計にする。
    """
    try:
        # 派生ページへ移動
        page.goto(derived_url, timeout=15000)

        # 最低限DOM構築完了を待機
        page.wait_for_load_state("domcontentloaded")

        # 第3回で作った安全取得関数を再利用
        return safe_text_content(page, sel_derived_tag1)
    except Exception:
        # 派生ページだけ失敗しても全体停止しない
        return None

4. main.py で一次データと派生データを統合する

ここでは、records をループしながら派生情報を追記します。

from extractor import (
    build_derived_url,
    fetch_specific_tag1_from_derived_page,
)

# ...(第3回までの処理で records を作成済みという前提)

for record in records:
    # 商品ID1から派生URLを作る
    derived_url = build_derived_url(settings.detail_base_url, record.product_id1)
    record.derived_url = derived_url

    # product_id1 が空のケースは、派生取得をスキップ
    if not derived_url:
        continue

    # 派生ページへアクセスして specific_tag1 を取得
    specific_tag1 = fetch_specific_tag1_from_derived_page(
        page=page,
        derived_url=derived_url,
        sel_derived_tag1=settings.sel_derived_tag1,
    )
    record.specific_tag1 = specific_tag1

この時点で、records の各要素に以下がそろいます。

  • 一次ページ由来: product_id1, product_id2, product_name, issue1, issue2
  • 派生ページ由来: derived_url, specific_tag1

5. 実行時の確認ポイント

  • product_id1 がある行に derived_url が入っている
  • 派生ページに値がある行だけ specific_tag1 が埋まる
  • 派生取得失敗があっても、一次取得済みデータは保持される

確認用の簡易表示例:

for rec in records[:5]:
    print(
        rec.product_id1,
        rec.product_name,
        rec.issue1,
        rec.issue2,
        rec.specific_tag1,
    )

ハマりどころと対策

product_id1 が空で派生URLが作れない

  • 現実運用ではよく起きます。
  • 本文コードでは build_derived_url() で None を返し、スキップ継続します。

派生ページだけセレクタが違う

  • 一次ページと混ぜず、SEL_DERIVED_TAG1 を独立して管理しましょう。
  • 将来的には selectors_derived.py のように分離してもよいです。

連続アクセスで不安定になる

  • 必要に応じて wait_for_timeout(300~1000ms) を入れて負荷を下げます。
  • 次回、失敗URL保存とリトライ設計を入れて運用安定化します。

まとめ

第4回では、派生URL処理まで実装できました。

  • product_id1 から派生URLを生成
  • 派生ページから specific_tag1 を取得
  • 一次抽出データへ統合

次回(第5回)は、この統合済みデータを最終フォーマットでテキスト出力し、
失敗URL記録・リトライ方針まで含めて運用しやすい形に仕上げます。

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?