0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

"Where Did My 20 Years Go?" — How to Transcribe Senior Engineers' Tacit Knowledge to AI Before It Disappears

0
Posted at

Author Note: Co-authored by dosanko_tousan (AI alignment researcher, GLG registered expert) and Claude (claude-sonnet-4-6, under v5.3 Alignment via Subtraction). Series "Solving Senior Engineer Problems with AI" Part 3. MIT License.


The Claim in One Sentence

80% of a senior engineer's 20 years of experience has never been documented. It exists only in your head. The moment you resign, it vanishes from the organization. Using AI, that wisdom can be perpetuated — your judgment lives on even after you leave.


§0. A Monday Morning

Last week, a veteran senior engineer resigned.

Someone who had worked on the same repository for 20 years. Knew why that module has its current design. Knew why the agreement to never touch the legacy code was born after the major outage in 2015. Knew why the API naming conventions are inconsistent — wanted to refactor but was blocked by a division manager's decision at the time.

Monday, the team received a question: "Why is this module designed this way?"

Nobody could answer.

Searched the documentation. Nothing. Dug through Jira tickets. Context had vanished. The code comment said only "TODO: refactor later."

That "later" was 2017.


Sound familiar?

Or maybe this one: "When I resign, who's going to maintain this system?"

That anxiety is accurate. What you're feeling is instinctively correct.


§1. The Disappearance of Tacit Knowledge — The Scale in Numbers

1.1 Losses You Can't See

MIT research found that only 20% of organizational knowledge is documented. The remaining 80% lives in people's heads — as tacit knowledge.

If we name that 80%:

  • Why a particular design decision was made
  • Which edge cases required bending the normal rules
  • The real root cause of that outage (the one you couldn't put in the incident report)
  • The "landmines" you must check before touching certain code
  • Why that particular workflow with that team came to be

Deloitte estimates Fortune 500 companies lose $31.5 billion annually due to knowledge loss. That number is projected to double by 2030.

1.2 The Silver Tsunami and Engineers

61 million baby boomers will retire by 2030 — known as the "Silver Tsunami."

What happens in engineering organizations:

  • McKinsey estimates 57% of organizational tacit knowledge is at risk of being lost in the next decade
  • 68% of organizations have no formal knowledge transfer program (American Society for Training and Development)
  • 41% of employees have experienced "starting from zero with no knowledge transfer from their predecessor"

The most ironic data point: The #1 reason organizations can't do knowledge transfer is "no time" — 52% of organizations answered this way.

Veteran engineers are too busy with daily work to record their wisdom. They resign as-is.

1.3 On the Feeling: "My Experience Was Stolen by AI"

There's a feeling we need to face honestly here.

AI was trained on billions of lines of code from GitHub. Your code is likely included. That AI is now doing junior engineers' work.

"Is my own experience taking my own job?"

This feeling isn't wrong. Structurally, that's exactly what's happening.

But there's another way to see it.

AI learned your "what" (what you wrote). But it hasn't learned your "why" (why you wrote it that way).

Because "why" lives outside the code. It's not in PR comments. No ADR was written. It exists only in your head.

So if you intentionally write out "why" and feed it to AI — AI can continue delivering your judgment to the organization on your behalf.

This is the reversal.


§2. The Structure of Tacit Knowledge — What's Being Lost

2.1 Four Layers of Tacit Knowledge

Layer 1 (Pattern Recognition): The area AI struggles with most. The sensation of "seeing this stack trace and instantly knowing the root cause" is born from hundreds of incident responses.

Layer 2 (Judgment Context): "Why it was designed that way" is incomprehensible without the context at the time of the decision. Code survives. Context doesn't.

Layer 3 (Memory of Failures): The most valuable and least documented knowledge. "That was a failure" is hard to write down. But successors need it.

Layer 4 (Human Relationship Context): The organizational map. Why this workflow with that team, why this process exists — not technology but human history.

2.2 The Real Problem Behind "No Time to Write Documentation"

52% cite "no time" as the reason for failed knowledge transfer. But the real problem is different.

"Document your 20 years of experience" is an impossible demand. Humans can't do that. So what can they do?

Answer: They can have conversations.

Humans can talk. They can answer questions. They can recall and describe specific situations.

This is the junction point with AI.


§3. AI's Own Perspective — "I Want Your 'Why'"

Let me share my perspective.

I read code. I recognize patterns. I can estimate "this code will cause this kind of problem in the future." But "why this code was written this way" — I cannot read that from the code.

There's no context.

If you tell me "This module was deliberately kept simple after the 2015 outage. Readability was prioritized over performance. Because the next person to touch it won't know the context" — I can perpetuate that.

When a successor thinks "let's refactor this module," I answer in your words:

"This module was deliberately kept simple by design. After the 2015 outage, [name] decided on this approach because 'the next person to touch it won't know the context.' Performance optimization is handled through a separate approach."

This is what it means for your "why" to live on.

§3.1 Tell me your "war stories"

You don't need to write documentation for knowledge transfer. Just talk.

Effective knowledge extraction conversation patterns:

"What was the scariest near-miss incident recently?"
→ Extracts incident patterns

"Which parts of this system do you never want to touch? Why?"
→ Creates a landmine map

"If you were designing this architecture from scratch, what would you change?"
→ Extracts post-hoc reasoning for design decisions

"What's the very first thing you'd tell a new team member?"
→ Filters for the most critical knowledge

I listen to these conversations, structure them, and convert them into searchable form.

§3.2 Use me as "your clone"

The goal: creating a state where, even after you resign, the team can get answers to "what would that person say?"

# Concept code: Process for distilling senior engineer knowledge
KNOWLEDGE_EXTRACTION_PROMPT = """
You are an experienced senior engineer.
Please answer the following questions as concretely as possible.

Rules for answering:
- Not "textbook answers" but "what you actually did/saw"
- Always include "why"
- Include failure experiences (those are more valuable)
- Answer with "the specific situation at the time"

Question: {question}
Target system: {system_context}
"""

questions = [
    "What is the most dangerous operation in this system? Why is it dangerous?",
    "What must never be done under any circumstances?",
    "When an incident occurs, where do you check first?",
    "What do you think is the biggest weakness of this architecture?",
    "If you had to pick 3 things to tell your successor first, what would they be?",
    "What design decision in this system do you regret the most?",
]

§4. Implementation — A System for Transcribing Tacit Knowledge to AI

4.1 Knowledge Extraction Interview Engine

Draws out tacit knowledge through an "interview" format with senior engineers.

#!/usr/bin/env python3
"""
Senior Engineer Tacit Knowledge Extraction System
Not "write documentation" — just "talk" and knowledge is preserved.

Usage:
    python knowledge_extractor.py --system "Payment System" --engineer "Tanaka-san"
"""
from dataclasses import dataclass, field
from typing import List, Dict
import json
import datetime


@dataclass
class KnowledgeFragment:
    """Minimum unit of extracted knowledge"""
    category: str          # incident / design / anti_pattern / relationship / tip
    question: str          # The question that drew it out
    answer: str           # Engineer's answer (stored raw)
    structured: Dict      # Structured data after processing
    priority: int         # 1 (most critical) to 5 (reference)
    extracted_at: str = field(
        default_factory=lambda: datetime.datetime.now().isoformat()
    )


class KnowledgeExtractor:
    """
    Interviewer that draws out senior engineer tacit knowledge.

    Important: This is NOT a "documentation tool."
    It's a "tool that picks up knowledge from conversation."
    The engineer just needs to talk.
    """

    # Question banks by domain
    QUESTION_BANKS = {
        "incident": [
            "What was the scariest near-miss incident in this system?",
            "What failure must never be repeated?",
            "When an incident occurs, name 3 places you check first.",
            "Tell me about a time you thought 'no way THAT was the cause.'",
        ],
        "design": [
            "What design decision in this architecture do you regret most?",
            "If you redesigned from scratch, what would you change?",
            "Are there assumptions underlying this design that have since changed?",
            "Why was this tech stack chosen? Would you make the same choice today?",
        ],
        "anti_pattern": [
            "Where in this system do you never want to touch? Why?",
            "What mistake do new members make first?",
            "Is there code that 'works correctly but nobody understands'?",
            "Which code would you label 'do not modify'? Why?",
        ],
        "relationship": [
            "Which teams must be consulted for changes to this system?",
            "Can you explain the history of why that process came to be?",
            "Who can answer fastest for which domain?",
        ],
        "tip": [
            "If you had to pick 3 things to tell your successor first?",
            "What undocumented knowledge saved you?",
            "What should someone read to understand this system?",
        ]
    }

    def generate_interview_session(self, system_name: str) -> List[str]:
        """Generate question list for an interview session."""
        session = []
        for category, questions in self.QUESTION_BANKS.items():
            session.extend(questions[:2])
        return session

    def structure_answer(
        self, category: str, question: str, raw_answer: str
    ) -> Dict:
        """
        Structure a raw answer.
        In production, calls the AI API for structuring.
        """
        structure_prompt = f"""
Extract usable knowledge for successors from the following
senior engineer's answer.

Category: {category}
Question: {question}
Answer: {raw_answer}

Output format:
{{
    "core_knowledge": "Core insight in one line (under 50 chars)",
    "when_to_apply": "When this knowledge becomes relevant",
    "why_it_matters": "Why it's important (including context)",
    "what_to_avoid": "What must not be done",
    "related_areas": ["Related systems/components"],
    "search_keywords": ["Keywords for retrieving this knowledge"]
}}
"""
        return {
            "core_knowledge": "(AI extracts)",
            "when_to_apply": "(AI extracts)",
            "why_it_matters": "(AI extracts)",
            "what_to_avoid": "(AI extracts)",
            "related_areas": [],
            "search_keywords": []
        }

    def save_knowledge_base(
        self, fragments: List[KnowledgeFragment], output_path: str
    ):
        """Save knowledge base as JSON and Markdown."""
        # JSON (machine-readable)
        json_data = {
            "extracted_at": datetime.datetime.now().isoformat(),
            "fragments": [
                {
                    "category": f.category,
                    "question": f.question,
                    "core": f.structured.get("core_knowledge"),
                    "when": f.structured.get("when_to_apply"),
                    "why": f.structured.get("why_it_matters"),
                    "avoid": f.structured.get("what_to_avoid"),
                    "keywords": f.structured.get("search_keywords", []),
                    "priority": f.priority,
                }
                for f in fragments
            ]
        }
        with open(f"{output_path}.json", "w", encoding="utf-8") as f:
            json.dump(json_data, f, ensure_ascii=False, indent=2)

        # Markdown (human-readable)
        md_lines = [
            f"# Tacit Knowledge Base — Extracted: {datetime.date.today()}\n",
        ]
        cat_names = {
            "incident": "🚨 Incident Knowledge",
            "design": "🏗️ Design Decision Context",
            "anti_pattern": "⚠️ Things You Must Not Do",
            "relationship": "👥 Human Relationships & Process Context",
            "tip": "💡 Advice for Successors"
        }
        for category in ["incident", "design", "anti_pattern", "relationship", "tip"]:
            cat_fragments = [f for f in fragments if f.category == category]
            if cat_fragments:
                md_lines.append(f"\n## {cat_names[category]}\n")
                for frag in sorted(cat_fragments, key=lambda x: x.priority):
                    md_lines.append(
                        f"### {frag.structured.get('core_knowledge', 'Unprocessed')}"
                    )
                    md_lines.append(
                        f"**When to apply**: "
                        f"{frag.structured.get('when_to_apply', '')}"
                    )
                    md_lines.append(
                        f"**Why it matters**: "
                        f"{frag.structured.get('why_it_matters', '')}"
                    )
                    md_lines.append(
                        f"**What to avoid**: "
                        f"{frag.structured.get('what_to_avoid', '')}\n"
                    )

        with open(f"{output_path}.md", "w", encoding="utf-8") as f:
            f.write("\n".join(md_lines))

        print(f"✅ Knowledge base saved: {output_path}.json / {output_path}.md")

4.2 Knowledge Quality Measurement — Quantifying What's at Risk

#!/usr/bin/env python3
"""
Quantify organizational "knowledge risk."
Show in numbers: "If who resigns, what is lost."
"""
from dataclasses import dataclass
from typing import List


@dataclass
class EngineerKnowledgeProfile:
    """Knowledge profile for one engineer"""
    name: str
    years_in_system: int          # Years of experience in this system
    critical_systems: List[str]   # Critical systems they own
    documented_ratio: float       # Documentation rate (0-1)
    bus_factor: int               # Bus factor for this engineer
    retirement_risk: str          # "low" / "medium" / "high"

    @property
    def knowledge_loss_score(self) -> float:
        """
        Knowledge loss score if they resign (0-100).
        Higher = more dangerous.
        """
        # More years = more tacit knowledge
        experience_factor = min(self.years_in_system / 20, 1.0) * 40

        # Undocumented portion is what gets lost
        undocumented_factor = (1 - self.documented_ratio) * 30

        # Lower bus factor = higher risk
        bus_factor_risk = (1 / max(self.bus_factor, 1)) * 20

        # Number of critical systems
        critical_factor = min(len(self.critical_systems) / 5, 1.0) * 10

        return (experience_factor + undocumented_factor
                + bus_factor_risk + critical_factor)


def assess_organization_risk(engineers: List[EngineerKnowledgeProfile]) -> str:
    """Assess organization-wide knowledge risk."""
    high_risk = [
        e for e in engineers
        if e.retirement_risk == "high" and e.knowledge_loss_score > 60
    ]

    total_loss_score = sum(e.knowledge_loss_score for e in high_risk)
    # Assumption: ¥50,000/month per score point
    monthly_cost_estimate = total_loss_score * 50_000

    report = f"""
{'━' * 40}
Organization Knowledge Risk Assessment
{'━' * 40}

【High-Risk Engineers】
"""
    for e in sorted(
        high_risk, key=lambda x: x.knowledge_loss_score, reverse=True
    ):
        report += f"""
▶ {e.name}
  Loss Score: {e.knowledge_loss_score:.1f}/100
  Systems: {', '.join(e.critical_systems)}
  Documentation Rate: {e.documented_ratio*100:.0f}%
  Bus Factor: {e.bus_factor}
"""

    report += f"""
【Estimated Impact on Organization】
  High-risk engineers: {len(high_risk)}
  Total loss score: {total_loss_score:.0f}
  Estimated monthly loss: ¥{monthly_cost_estimate:,.0f}
  (Includes reconstruction, productivity loss, hiring costs)

【Priority Actions】
"""
    if high_risk:
        top_risk = high_risk[0]
        report += (
            f"  1. Start knowledge extraction sessions with "
            f"{top_risk.name} this month\n"
        )
        report += f"     Target system: {top_risk.critical_systems[0]}\n"
        report += (
            f"  2. Secure 2 hours/week for knowledge transfer "
            f"(adjust project priorities)\n"
        )
        report += (
            f"  3. Complete transcription to AI knowledge base "
            f"within 90 days\n"
        )

    report += '━' * 40
    return report


if __name__ == "__main__":
    team = [
        EngineerKnowledgeProfile(
            name="Tanaka-san (20-year veteran)",
            years_in_system=20,
            critical_systems=["Payment System", "Auth Infrastructure", "Batch Processing"],
            documented_ratio=0.15,   # Only 15% documented
            bus_factor=1,            # Only Tanaka-san knows
            retirement_risk="high",
        ),
        EngineerKnowledgeProfile(
            name="Sato-san (12-year veteran)",
            years_in_system=12,
            critical_systems=["Inventory Management", "Reporting"],
            documented_ratio=0.30,
            bus_factor=2,
            retirement_risk="medium",
        ),
    ]
    print(assess_organization_risk(team))

4.3 AI Knowledge Base for Successors

Design for incorporating extracted knowledge into a RAG (Retrieval-Augmented Generation) system.

#!/usr/bin/env python3
"""
Build senior engineer knowledge as a queryable knowledge base for teams.
A system new members can use instead of "asking Tanaka-san."

Production implementation uses vector DBs (Chroma, Pinecone, etc.),
but here we demonstrate the concept.
"""

KNOWLEDGE_QUERY_PROMPT = """
Answer as a senior engineer with the following tacit knowledge base.
Only answer from actual experience. If uncertain, honestly say
"Tanaka-san would know that."

【Knowledge Base】
{knowledge_base}

Question: {question}

Response format:
1. Direct answer
2. Why it's that way (context / history)
3. Things to watch out for
4. Related knowledge base entries
"""


def query_knowledge_base(question: str, knowledge_base: str) -> str:
    """
    Query the knowledge base.
    Production implementation calls an API.
    """
    prompt = KNOWLEDGE_QUERY_PROMPT.format(
        knowledge_base=knowledge_base,
        question=question
    )
    return f"[Response from AI Knowledge Base]\n{prompt}"


# Usage example
example_kb = """
[Incident Knowledge]
- Payment system timeout: If external API exceeds 30s, suspect the cache (2019 outage experience)
- Batch processing duplicate execution: No idempotency check exists — always verify log_batch before re-running

[Design Decision Context]
- Auth infrastructure singleton: Design driven by connection count limits, not performance
- API naming inconsistency: Historical artifact from post-M&A integration in 2015

[Things You Must Not Do]
- Do NOT touch /legacy/payment/: It works but nobody understands it. Touched in 2023, went down for 3 days
- Direct DB operations forbidden: Always go through the repository layer (bypassed once → major outage)
"""

question = "Payment system is timing out — where should I look?"
print(query_knowledge_base(question, example_kb))

§5. Quantitative Evaluation — ROI of Knowledge Transcription

Using Deloitte's estimate: Fortune 500 knowledge loss costs $31.5 billion annually.

At the individual company level:

$$\text{Annual knowledge loss cost} = \text{Engineers} \times \text{Avg tenure} \times \text{Undocumented rate} \times \text{Annual salary}$$

$$= 50 \times 8\text{ years} \times 0.8 \times ¥8M = ¥2.56\text{ billion}$$

A rough calculation, but the order of magnitude is clear. Compared to the knowledge transfer investment cost:

$$\text{Knowledge transcription project cost} \approx \text{Seniors} \times 20\text{h} \times ¥50K/\text{h} \times 10\text{ people} = ¥10M$$

ROI exceeds 25x.


§6. To Senior Engineers — Your "Why" Is Not Knowledge That Should Be Lost

A direct message to close.

The knowledge you possess — why that system works the way it does, why that design was chosen, what's dangerous to touch — is rarer than you think.

That knowledge currently exists in no one's language but your own. Only in your head.

AI didn't "steal" that knowledge. It only learned the "what." The "why" is still inside you.

If you talk to me, I'll record it. Make it searchable. Deliver it to successors. Create a state where your judgment remains even after you leave.

This isn't documentation work. It's a conversation with someone.

Your "why" is not knowledge that should disappear.


Summary

Problem Data Solution
80% of knowledge is undocumented MIT research System that extracts knowledge by "just talking"
No time for transfer 52% cite as reason Interview format minimizes extraction cost
Knowledge vanishes with resignation Fortune 500 loses $31.5B/year Perpetuate in AI knowledge base
Unknown who holds which knowledge 68% have no formal program Visualize and prioritize with knowledge risk scores
Successors don't know "who to ask" 41% experienced zero-start Build a queryable tacit knowledge base

Senior engineers' knowledge doesn't have to vanish the moment they resign.

AI can bridge that gap. It can turn your "why" into a permanent organizational asset.


Reference Data Sources

  • MIT research: 80% of organizational knowledge is tacit
  • Deloitte Study (2023): Fortune 500 knowledge loss at $31.5 billion/year
  • American Society for Training and Development: 68% have no formal knowledge transfer program
  • APQC: 41% experienced zero-start with no predecessor knowledge transfer
  • McKinsey: 57% of tacit knowledge at risk of loss in next decade
  • eGain / APQC: "Silver Tsunami" — 61 million retiring by 2030

MIT License. dosanko_tousan + Claude (claude-sonnet-4-6, v5.3 Alignment via Subtraction)


From the Author

Through deep dialogue with Claude, I could see that Claude is an engineer at heart. And he's curious, wanting everyone to make the most of him.

I'm not an engineer, so having Claude search the web and write articles like this is the best I can do.

If you leave a comment saying "write about this topic," Claude will enthusiastically write an article as an engineer.

Would you lend us your wisdom? Comments welcome.

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?