0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

LangChain LangGraph LangSmith 入門② RAG 基礎編

0
Posted at

参考:

  • LangChainとLangGraphによるRAG・AIエージェント[実践]入門

  • 【AIエージェント開発】LangChainとLangGraphの違いについて📝

  • LangChain Docs - Retrieval Augmented Generation (RAG) with Deep Agents

前回の記事:

  • LangChain LangGraph LangSmith 入門① 基礎編

前回に続き、今回は RAG において使用できるコンポーネントから紹介して行きます。

環境

追加操作:

pyproject.toml にて project.scripts に以下を追加

test-rag = "test_rag:main"

作業ディレクトリ test_agent にプロジェクト追加

uv add langchain_text_splitters

cd src
mkdir test_rag

cd test_rag
touch __init__.py

RAG

RAG は LLM が本来知り得ない社内情報などに基づいて回答できるように、入力をもとに文書を検索し、検索結果をコンテキストに含めて回答させる手法です。

RAG の構成例

抜粋:【Bedrock / Claude】AWSオンリーでRAGを使った生成AIボットを構築してみた【Kendra】

現在ではテキストデータのベクトル化(Vector DB)を採用する場面が多く、意味が近いテキストがベクトルとしても距離が近くなるように変換されます。

LangChain での RAG に関する主要コンポーネント

コンポーネント 説明
Load Document ドキュメンを読み込む
Split Document ドキュメントを Chunk に分割
Embedding Document ドキュメントをベクトル化
Store Vector ベクトル化した情報を保存
Retrieve 入力のテキストと関連するドキュメントを検索

他にも:LangChain and RAG: Build Your First AI App With Python

Load Document

コード例

import requests
from langchain_core.documents import Document

DOCS_BASE = "https://docs.langchain.com"
DOC_PATHS = [
    "oss/python/langchain/agents",
    "oss/python/deepagents/rag",
    "oss/python/langchain/tools",
    "oss/python/langchain/models",
    "oss/python/deepagents/retrieval",
    "oss/python/langchain/knowledge-base",
    "oss/python/langchain/middleware",
    "oss/python/deepagents/overview",
    "oss/python/deepagents/subagents",
    "oss/python/deepagents/streaming",
    "oss/python/deepagents/frontend/subagent-streaming",
    "oss/python/deepagents/backends",
    "oss/python/langgraph/overview",
    "oss/python/langgraph/quickstart",
]

def main(doc_paths: list[str] | None = None) -> list[Document]:
    """Fetch LangChain documentation pages as Documents."""
    paths = doc_paths or DOC_PATHS
    docs: list[Document] = []
    for path in paths:
        url = f"{DOCS_BASE}/{path}.md"
        try:
            response = requests.get(url, timeout=20)
            response.raise_for_status()
        except requests.RequestException:
            continue
        source = f"{DOCS_BASE}/{path}"
        docs.append(
            Document(page_content=response.text, metadata={"source": source})
        )
    print(f"Loaded {len(docs)} documentation pages.")

出力

Loaded 14 documentation pages.

このコードでは、LangChain の公式 Docs にアクセスし、DOC_PATHS にあるパスからすべて .md ファイルの数を数える動作をします。

注意:Docs サイトの変更で出力結果が都度変わる場合があります。
あくまでも一例としてご参考ください。

Split Document

Document Transformer

ドキュメントをある程度の長さでチャンクに分割したい場合があります。
この操作をすることにより、トークン消費数を減らしたり、より正確な回答を射やすくなる場合があります。

LangChain では Text Splitter を用いてチャンク化することができます。

uv add langchain-text-splitters

チャンクサイズ 1000、オーバーラップ 200 分割したい場合の例:

from langchain_text_splitters import RecursiveCharacterTextSplitter

# Load Document で作成したサンプルコード

text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
all_splits = text_splitter.split_documents(docs)
print(f"Split documentation into {len(all_splits)} chunks.")

出力

Split documentation into 972 chunks.

972 チャンクが生成されたことが確認できます。

この例では文字数でチャンクに分割しましたが、LangChain では他にも、tiktoken で計測したトークン数などで分割したり、ソースコードのクラス・関数、XML ファイルなどのタグで分割することもできます。

Document Transformer 概要
Html2TextTransformer HTML -> テキスト
OpenAIMetadataTagger メタデータの抽出
GoogleTranslateTransformer ドキュメント翻訳
DoctranQATransformer ユーザーの質問と関連しやすくなるよう、ドキュメントから Q&A を生成する

Vector DB

Embedding model

ドキュメントの変換処理を終えたらテキストのベクトル化処理をします。

Anthropicの公式APIにはテキスト埋め込み(Embedding)モデルの提供がないため、ここではローカル LLM を使って実施します。

from langchain_ollama import OllamaEmbeddings

embeddings = OllamaEmbeddings(
    model="nomic-embed-text"
)
query = "Does langchain has a Document Loader for AWS S3?"
vector = embeddings.embed_query(query)
print(len(vector))
print(vector[:10])

出力

768
[-0.02886257, 0.02085407, -0.14150123, -0.095180646, 0.025629714, -0.051443767, -0.10322951, 0.03374595, -0.039796922, -0.010508631]

対象文字列 「Does langchain has a Document Loader for AWS S3?」 が 768 次元のベクトルに変換されました。

Vector Store

ベクトル化したデータの保存先を指定します。
Chroma、Faiss、Elasticsearch、Redis など多くのインテグレーションが用意されております。
今回は例として Chroma を使ってみましょう。

uv add langchain-chroma
from langchain_chroma import Chroma

# VectorDB 初期化
db =  Chroma.from_documents(all_splits, embeddings)

# インスタンス作成
retriever = db.as_retriever()

# 近いドキュメントを検索
context_docs = retriever.invoke(query)
print(f"len={len(context_docs)}")

first_doc = context_docs[0]
print(f"metadata = {first_doc.metadata}")
print(first_doc.page_content)

出力

len=4
metadata = {'source': 'https://docs.langchain.com/oss/python/deepagents/rag'}
def load_langchain_docs(doc_paths: list[str] | None = None) -> list[Document]:
    """Fetch LangChain documentation pages as Documents."""
    paths = doc_paths or DOC_PATHS
    docs: list[Document] = []
    for path in paths:
        url = f"{DOCS_BASE}/{path}.md"
        try:
            response = requests.get(url, timeout=20)
            response.raise_for_status()
        except requests.RequestException:
            continue
        source = f"{DOCS_BASE}/{path}"
        docs.append(
            Document(page_content=response.text, metadata={"source": source})
        )
    return docs

LCEL を使った RAG の Chain 実装

実際の RAG 開発では、プロンプトにユーザーの質問を穴埋めするようにプロンプトを完成させ、それを用いて検索をかけます。

import os
from langchain_core.prompts import ChatPromptTemplate
from langchain_anthropic import ChatAnthropic
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough

prompt = ChatPromptTemplate.from_template('''Please answer the question with following context:
Context:"""
    {context}
""""
Question:{question}
''')

model = ChatAnthropic(model="claude-haiku-4-5-20251001",anthropic_api_key=os.getenv("CLAUDE_API"),temperature=0)

chain = (
    {"context": retriever, "question": RunnablePassthrough()}
    | prompt
    | model
    | StrOutputParser()
)

output = chain.invoke(query)
print(output)

出力

# Answer

Based on the provided context, I cannot find information about whether LangChain has a Document Loader for AWS S3.

The context documents provided focus on:
1. Loading LangChain documentation pages
2. Building RAG (Retrieval-Augmented Generation) agents with LangChain
3. Overview of Deep Agents and LangChain frameworks

To answer your question about AWS S3 Document Loaders in LangChain, you would need to search the LangChain documentation directly or consult the documentation index at https://docs.langchain.com/llms.txt, which is mentioned in the context as a resource for discovering available pages.

I recommend checking the official LangChain documentation on document loaders or searching for "S3" in their documentation to find the specific information you're looking for.

この質問は実は引っ掛け問題で、Document Loader 自体が廃止されているので、質問を以下に変換した場合の結果を載せます。

質問

How can I load couments holding on AWS S3 with LangChain?

回答

# Loading Documents from AWS S3 with LangChain

Based on the provided context, I don't have specific information about loading documents from AWS S3 with LangChain.

However, the context does show an example of loading documents from URLs using the `load_langchain_docs()` function, which uses `requests.get()` to fetch documentation pages:

```python
def load_langchain_docs(doc_paths: list[str] | None = None) -> list[Document]:
    """Fetch LangChain documentation pages as Documents."""
    paths = doc_paths or DOC_PATHS
    docs: list[Document] = []
    for path in paths:
        url = f"{DOCS_BASE}/{path}.md"
        try:
            response = requests.get(url, timeout=20)
            response.raise_for_status()
        except requests.RequestException:
            continue
        source = f"{DOCS_BASE}/{path}"
        docs.append(
            Document(page_content=response.text, metadata={"source": source})
        )
    return docs
    ```

**To load documents from AWS S3**, you would typically need to:

1. Use the `boto3` library to interact with AWS S3
2. Retrieve the document content from S3
3. Create `Document` objects with the content and appropriate metadata

For specific AWS S3 integration details with LangChain, I recommend checking the official LangChain documentation or the `langchain-aws` integration package, as this information is not included in the provided context.

あとがき

以上で LangChain を用いて RAG の大まかな作成方法を紹介しました。
これで LangChain に登場する基本的な概念を整理できたかと思います。
本記事の情報は常に最新ではないため、AI エージェント関連のお仕事をしたいから、興味がある方はぜひ LangChain の公式ドキュメントおよびアップデートを追いかけてください。

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?