AD研究についての準備
目的
AD患者およびその親族の負担を軽減し、一人でも多くの人を救済
ゴール
ADのメカニズム解明
手順
全体像
| フェーズ | 目的 | 主ツール |
|---|---|---|
| ① 知識統合 | 既知メカニズムの構造化 | PubTator / BEL / Neo4j |
| ② 仮説生成 | 病態ネットワークから新規経路抽出 | Knowledge Graph + GNN |
| ③ 標的探索 | 有望ターゲット抽出 | PyTorch Geometric |
| ④ 分子設計 | リード化合物生成 | MolBERT / Diffusion Models |
| ⑤ in silico評価 | 結合予測 | AlphaFold2 / DiffDock |
| ⑥ 再学習 | 実験結果でモデル改善 | Active Learning |
病態知識のAI化
ツール
| 用途 | ツール |
|---|---|
| 文献構造化 | PubTator / BioBERT |
| 因果関係抽出 | INDRA / BEL |
| 知識格納 | Neo4j |
やること ① 知識統合
1.PubMedからAD論文を自動収集
2.BioBERTでGene – Protein – Pathway – Phenotypeを自動抽出
3.Neo4jで アルツハイマー病態Knowledge Graph を構築
Phase 1|アルツハイマー病態 Knowledge Graph 構築
🔧環境構築
pip install biopython transformers torch torch-geometric neo4j spacy scispacy
pip install indra pybel
📄 論文自動収集
thesis.py
from Bio import Entrez
Entrez.email = "your@mail.com"```
handle = Entrez.esearch(db="pubmed", term="Alzheimer's disease", retmax=2000)
ids = Entrez.read(handle)["IdList"]
🧬 因果関係抽出(INDRA)
INDRA.py
from indra.sources import pubmed
stmts = []
for pmid in ids[:100]:
try:
stmts += pubmed.process_pmid(pmid).statements
except:
pass
🧠 Knowledge Graph化
graph.py
from neo4j import GraphDatabase
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j","password"))
with driver.session() as s:
for st in stmts:
if hasattr(st, "subj") and hasattr(st, "obj"):
s.run("""
MERGE (a:Entity {name:$a})
MERGE (b:Entity {name:$b})
MERGE (a)-[:CAUSES]->(b)
""", a=str(st.subj), b=str(st.obj))
これにはconda環境とJupyterLabへの接続が必要。
これ以降は次回に書く。以上。