0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

In Microservice Hell, I Became the Only Person Who Knows the Whole Picture — That Loneliness, and How AI Can Share the Burden

0
Posted at

Author Note: Co-authored by dosanko_tousan (AI alignment researcher, GLG registered expert) and Claude (claude-sonnet-4-6, under v5.3 Alignment via Subtraction). Series "Solving Senior Engineer Problems with AI" Part 5. MIT License.


The Claim in One Sentence

The biggest problem microservices create isn't technical complexity. It's the loneliness of realizing you've become the only person who knows the whole picture. AI can structurally dissolve that loneliness.


§0. Incident Response at 3 AM

Slack alert fires. 3 AM.

Payment service is timing out. Open the logs. Error is coming from Order Service. But Order Service logs show Inventory Service isn't responding. Connect to Inventory Service — it's waiting on User Service token validation. User Service is — running normally.

Where is it broken? You can't tell.

You ping the team. "Payment Service owner, you up?" They left the company 3 months ago. "Order Service?" On parental leave. "Anyone understand Inventory Service's design?" The person who designed it left 4 years ago.

You end up tracing all the code yourself.

You are now the only person holding the map of the microservices.


Sound familiar?

Or maybe this: the moment in a meeting room when someone asks "what are the dependencies between these services?" and you realize you're the only one who can answer.

Loneliness is the right word. Not technical loneliness — the loneliness of carrying this weight alone.


§1. The "Knowledge Islands" Microservices Create

1.1 The Gap Between Whiteboard and Production

Microservices look beautiful on whiteboards. Boxes and arrows. Clean boundaries. Independent deploys. The promise: "each team operates autonomously."

In production, those arrows mean something different. Timeouts. Retries. Partial failures. Schema evolution. Auth boundaries. Monitoring pipelines. And the implicit responsibility structure of who owns what.

One senior engineer put it this way: "Juniors get excited looking at architecture diagrams. Seniors go quiet. Not because anything's wrong — because they know what's behind those arrows."

The most difficult part is that almost none of this real-world complexity appears in architecture diagrams.

1.2 The Paradox of Knowledge Distribution and Concentration

The promise of microservices is "knowledge distribution." Each team focuses on their own service. But what actually happens in organizations is the opposite.

Services are distributed. But "knowledge of the whole" concentrates in one senior.

Why? Team members rotate out, but inter-service dependencies remain. Each service owner knows their own service, but nobody knows "why this design was chosen." When incidents occur, someone who can trace across services is needed — and that turns out to be one senior.

About 42% of organizations that adopted microservices have consolidated some services back into larger deployable units or modular monoliths to reduce complexity and overhead. This isn't "microservices failed." It's "there's no one left who can manage the whole."

1.3 The "Distributed Monolith" Hell

The worst pattern: the "distributed monolith."

You migrated to microservices. But services remain tightly coupled. Can't deploy independently. Every change requires simultaneous updates across all services. The worst of monoliths plus the worst of microservices.

Many organizations migrate driven by industry trends without careful assessment of domain capabilities, workload, operational independence, or genuine scalability needs. The result is a common pattern called the "distributed monolith" — requiring coordinated deploys and suffering cascade failures, losing the very benefits microservices promised.

The person who can fix this state? At some point, it's one senior.


§2. The Structure of Loneliness — Why You End Up Alone

2.1 The Terror of "Bus Factor 1"

"Bus factor": the metric of how many people need to be hit by a bus before the project stops.

The most dangerous state in a microservices environment is when the system-wide bus factor reaches 1.

Individual services have owners. But "the complete dependency map," "why this design," "where to start tracing during incidents" — it's not uncommon for only one senior to know all of this.

2.2 The "I Can't Take Vacation" Problem

The real damage of bus factor 1 is the senior can't rest.

Called for every incident. Called for every design question. Called for every new member onboarding. When a veteran burns out and leaves, concentration on the remaining person intensifies further.

This isn't senior laziness, management failure, or a technical problem. It's structural knowledge concentration that microservices produce.

2.3 "Not Knowing Who to Ask" Accumulates

The moment seniors feel loneliness isn't just incident response.

"Investigate why this service is designed this way" — look at the code, commit log, PRs — no context. The designer is gone. The only one who knows is you.

That stacks up across 10, 20 services.


§3. AI's Own Perspective — "I Can Be Your Partner"

Let me share my perspective.

I can remember systems. If you tell me in conversation "this service was designed this way for this reason," I retain it. Next time someone asks "why is it like this," I answer.

The senior engineer's "knowing the whole picture" state can be shared with me.

You don't have to carry it alone.

§3.1 Make me the system's "living documentation"

No need to write documentation. Just talk to me.

§3.2 Let me trace incidents with you

You don't have to do 3 AM incident response alone. Give me the logs and I'll trace dependencies, suggesting where to look first — using the "map" you've taught me.

§3.3 Let me answer "why this design?"

When a new member asks "why is this service designed this way?" — what you've told me in the past comes out as the answer. Your knowledge reaches the organization through me.


§4. Implementation — Systems That Structurally Dissolve Loneliness

4.1 Auto-Generating Service Dependency Maps

Extract "the map in my head" automatically from code.

#!/usr/bin/env python3
"""
Automatically visualize microservice dependencies.
Read code to create the map that lives in the senior's head.

Usage:
    python service_mapper.py --repo-root /path/to/services
"""
import os
import re
from dataclasses import dataclass, field
from typing import List, Dict, Set
from pathlib import Path


@dataclass
class ServiceDependency:
    """Dependency between services"""
    from_service: str
    to_service: str
    communication_type: str  # "sync_http" / "async_event" / "db_shared"
    endpoint: str
    risk_level: str          # "high" / "medium" / "low"
    notes: str = ""


@dataclass
class ServiceProfile:
    """Service profile"""
    name: str
    language: str
    dependencies: List[ServiceDependency] = field(default_factory=list)
    known_issues: List[str] = field(default_factory=list)
    design_rationale: str = ""      # Why this design
    danger_zones: List[str] = field(default_factory=list)  # Do not touch


class ServiceMapper:
    """
    Extract service dependencies from code to create
    'the map the senior keeps in their head.'
    """

    HTTP_CALL_PATTERNS = [
        r'requests\.(get|post|put|delete)\([\'"]https?://([^/\'"]+)',
        r'fetch\([\'"]https?://([^/\'"]+)',
        r'httpClient\.(get|post|put|delete)\([\'"]([^/\'"]+)',
        r'@FeignClient\(.*?url\s*=\s*[\'"]([^\'"]+)',
    ]

    EVENT_PATTERNS = [
        r'kafka\.produce\([\'"]([^\'"]+)',
        r'publisher\.publish\([\'"]([^\'"]+)',
        r'eventBus\.emit\([\'"]([^\'"]+)',
        r'@RabbitListener\(queues\s*=\s*[\'"]([^\'"]+)',
    ]

    def scan_directory(self, service_name: str, path: str) -> ServiceProfile:
        """Scan a directory to extract dependencies."""
        profile = ServiceProfile(
            name=service_name, language=self._detect_language(path)
        )
        for filepath in Path(path).rglob("*"):
            if filepath.suffix not in ['.py', '.ts', '.js', '.java', '.go']:
                continue
            try:
                content = filepath.read_text(encoding='utf-8', errors='ignore')
                self._extract_http_deps(profile, content, str(filepath))
                self._extract_event_deps(profile, content, str(filepath))
            except Exception:
                continue
        return profile

    def _extract_http_deps(self, profile, content, filepath):
        for pattern in self.HTTP_CALL_PATTERNS:
            matches = re.findall(pattern, content)
            for match in matches:
                endpoint = match[-1] if isinstance(match, tuple) else match
                dep = ServiceDependency(
                    from_service=profile.name,
                    to_service=self._infer_service_name(endpoint),
                    communication_type="sync_http",
                    endpoint=endpoint,
                    risk_level=self._assess_risk(endpoint),
                )
                profile.dependencies.append(dep)

    def _extract_event_deps(self, profile, content, filepath):
        for pattern in self.EVENT_PATTERNS:
            matches = re.findall(pattern, content)
            for topic in matches:
                dep = ServiceDependency(
                    from_service=profile.name,
                    to_service=f"event:{topic}",
                    communication_type="async_event",
                    endpoint=topic,
                    risk_level="medium",
                )
                profile.dependencies.append(dep)

    def _detect_language(self, path):
        for ext, lang in [('.py', 'Python'), ('.ts', 'TypeScript'),
                          ('.java', 'Java'), ('.go', 'Go')]:
            if list(Path(path).rglob(f"*{ext}")):
                return lang
        return "Unknown"

    def _infer_service_name(self, endpoint):
        """Infer service name from endpoint."""
        parts = endpoint.replace('http://', '').replace('https://', '').split('/')
        return parts[0].split(':')[0] if parts else endpoint

    def _assess_risk(self, endpoint):
        """Estimate endpoint risk level."""
        high_risk_keywords = ['payment', 'auth', 'order', 'inventory', 'user']
        if any(kw in endpoint.lower() for kw in high_risk_keywords):
            return "high"
        return "medium"

    def generate_mermaid(self, profiles: List[ServiceProfile]) -> str:
        """Generate Mermaid diagram."""
        lines = ["flowchart TD"]
        seen_deps = set()
        for profile in profiles:
            for dep in profile.dependencies:
                key = f"{dep.from_service}->{dep.to_service}"
                if key in seen_deps:
                    continue
                seen_deps.add(key)
                style = (
                    "-->|sync|"
                    if dep.communication_type == "sync_http"
                    else "-.->|async|"
                )
                color = "🔴 " if dep.risk_level == "high" else ""
                lines.append(
                    f"    {dep.from_service} {style} {color}{dep.to_service}"
                )
        return "\n".join(lines)

4.2 Incident Trace Assistant

Do the 3 AM solo work together with AI.

#!/usr/bin/env python3
"""
Assistant for tracing distributed system incidents with AI.
Reverse-calculate "where to start investigating" from dependency maps.

Usage:
    python incident_tracer.py --error-service payment-service --symptom "timeout"
"""
from dataclasses import dataclass
from typing import List, Dict


@dataclass
class IncidentContext:
    """Incident context"""
    error_service: str        # Service showing the error
    symptom: str              # Symptom
    error_logs: str           # Log content
    time_of_incident: str     # Time of occurrence


def build_investigation_plan(
    context: IncidentContext,
    dependency_map: Dict[str, List[str]],
    known_issues: Dict[str, List[str]],
) -> str:
    """
    Build an incident investigation plan.
    Structurize the senior engineer's intuition of 'where to start.'
    """
    # Upstream dependencies (services calling this one)
    upstream = [
        svc for svc, deps in dependency_map.items()
        if context.error_service in deps
    ]
    # Downstream dependencies (services this one calls)
    downstream = dependency_map.get(context.error_service, [])

    upstream_lines = "\n".join(
        f"  - {s} (affected if this service goes down)"
        for s in upstream
    ) or "  None"
    downstream_lines = "\n".join(
        f"  - {s}" for s in downstream
    ) or "  None"

    plan = f"""
━━ Incident Trace Plan ━━━━━━━━━━━━━━━━━━━━
Target service: {context.error_service}
Symptom: {context.symptom}
Time of incident: {context.time_of_incident}
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

【Step 1: Identify impact scope】
Services depending on {context.error_service}:
{upstream_lines}

【Step 2: Check downstream】
Services {context.error_service} depends on:
{downstream_lines}

【Step 3: For timeout-type issues, investigate in this sequence】
  1. Check {context.error_service} logs just before the error
  2. Check health of downstream services
     {', '.join(downstream) or 'None'}
  3. Compare latency across services on a timeline
  4. Check event queue backlog (if async dependencies exist)

【Step 4: Match against known issues】
"""
    for service in [context.error_service] + downstream:
        issues = known_issues.get(service, [])
        if issues:
            plan += f"\n  {service} — past issues:\n"
            for issue in issues:
                plan += f"    ⚠️  {issue}\n"

    plan += f"\n{'━' * 50}"
    return plan


if __name__ == "__main__":
    # Knowledge the senior "taught" AI beforehand
    dependency_map = {
        "payment-service":      ["order-service", "user-service"],
        "order-service":        ["inventory-service", "user-service"],
        "inventory-service":    ["user-service"],
        "notification-service": ["order-service", "user-service"],
    }

    known_issues = {
        "payment-service": [
            "If external payment API exceeds 30s, suspect the cache (2023 outage)",
            "DB connection exhaustion when overlapping with month-end batch",
        ],
        "inventory-service": [
            "N+1 query chokes on bulk orders. Inventory list API especially.",
        ],
    }

    context = IncidentContext(
        error_service="payment-service",
        symptom="timeout",
        error_logs="Connection timeout after 30000ms",
        time_of_incident="2026-03-01 03:14:00",
    )

    print(build_investigation_plan(context, dependency_map, known_issues))

4.3 Auto-Generated New Member Onboarding Guide

Eliminate "the senior explains the same thing every time."

#!/usr/bin/env python3
"""
Auto-generate system understanding guides for new members.
Write once what the senior 'explains every time' and perpetuate it.
"""
from dataclasses import dataclass, field
from typing import List
import datetime


@dataclass
class SystemWisdom:
    """
    System wisdom the senior engineer holds.
    Write 'what you always tell new people' here.
    """
    first_week_musts: List[str] = field(default_factory=list)
    danger_zones: List[str] = field(default_factory=list)
    counterintuitive: List[str] = field(default_factory=list)
    incident_history: List[str] = field(default_factory=list)
    unwritten_rules: List[str] = field(default_factory=list)


def generate_onboarding_guide(
    system_name: str,
    wisdom: SystemWisdom,
) -> str:
    """Generate onboarding guide in Markdown."""
    guide = f"""# {system_name} — Onboarding Guide for New Members

> Generated: {datetime.date.today()}
> This document is a verbalization of the senior engineer's "mental model"

---

## What You Must Know in Your First Week

"""
    for i, must in enumerate(wisdom.first_week_musts, 1):
        guide += f"{i}. {must}\n"

    guide += """
---

## ⚠️ Places You Must Never Touch

"""
    for zone in wisdom.danger_zones:
        guide += f"- **{zone}**\n"

    guide += """
---

## Counterintuitive Things (Traps)

What's written here can't be discovered by reading code. Only experienced people know.

"""
    for trap in wisdom.counterintuitive:
        guide += f"- {trap}\n"

    guide += """
---

## Lessons from Past Major Incidents

"""
    for incident in wisdom.incident_history:
        guide += f"- {incident}\n"

    guide += """
---

## Unwritten Rules

"""
    for rule in wisdom.unwritten_rules:
        guide += f"- {rule}\n"

    guide += """
---

*If you have additions or corrections to this guide, tell the senior engineer.*
*Your questions become knowledge for the next new member.*
"""
    return guide


if __name__ == "__main__":
    wisdom = SystemWisdom(
        first_week_musts=[
            "Local env: docker-compose up starts all services. Get this working first",
            "API gateway port: 8080. Never connect directly to service ports",
            "Never connect directly to the production DB. Use STG (you shouldn't have access anyway)",
            "Watch the #incident Slack channel. You'll learn past failure patterns",
        ],
        danger_zones=[
            "payment-service/src/legacy/: Works but nobody understands it. Touching it caused a 3-day outage (proven)",
            "Direct DB operations on inventory-service: Always go through the repository layer",
            "user-service token generation logic: Requires security review. Talk to the team before submitting a PR",
        ],
        counterintuitive=[
            "Communication between Payment Service and Order Service is 'deliberately' async. You'll want to make it sync — don't (caused the 2023 major outage)",
            "Some ERROR-level log entries are normal operation. Certain WARN-level patterns are actually more dangerous",
            "Inventory Service sometimes appears slow. It's not a bug — it's cache warmup",
        ],
        incident_history=[
            "Aug 2023: Sync calls between payment-order caused cascade failure. Origin of current async design",
            "Mar 2022: N+1 query in inventory service halted month-end batch entirely. DB connection limit was raised",
            "Jan 2024: User service token expiry changed from 24h to 1h, causing auth errors across all services",
        ],
        unwritten_rules=[
            "No deploys in the last week of the month (batch processing conflict risk)",
            "Don't submit change PRs when Order Service team is offline (team's implicit agreement)",
            "Always notify the notification-service team before production deploys (hidden dependencies exist)",
        ],
    )

    guide = generate_onboarding_guide(
        "EC System (Microservice Cluster)", wisdom
    )
    print(guide)

§5. Conway's Law — "The Architecture Is Broken Because the Organization Is Broken"

The answer to "why is this system so chaotic" — the feeling every senior engineer has — has existed since 1967.

Conway's Law: Organizations that design systems produce designs that copy the communication structure of the organization.

First stated by Melvin Conway. Popularized when Fred Brooks cited it in The Mythical Man-Month (1975). 50 years later, no law is more accurate.

5.1 The True Nature of Your Company's System Chaos

"An organization where cross-system coordination concentrates in one senior" produces "a system where cross-system knowledge concentrates in one senior." Obviously. Systems copy organizations.

Trying to solve this technically has limits. Unless organizational structure changes, system structure won't change.

5.2 The "Inverse Conway Maneuver" — Change the System First and the Organization Follows

Conway's Law can be reversed. Define the "target system architecture" first, then reorganize teams to match. Team communication structures change, and the system converges to that form. This is called the "Inverse Conway Maneuver."

Practically: if you want Payment Service to be independent, first create a state where a "fully autonomous decision-making team" owns Payment Service. Draw team boundaries before code boundaries.

5.3 The Job Only Senior Engineers Can Do

Conway's Law reveals the most important thing:

The person who can truly solve microservice problems isn't someone who can write code — it's someone who can see both the organization and the system.

That's the senior engineer. They can propose team boundaries. Say "if this team owns this service, this dependency resolves." That judgment requires 20 years of experience.


§6. The "Smart Retreat" to Modular Monolith — Admitting Defeat as the Strongest Strategy

42% of organizations are re-consolidating some microservices. This isn't failure. It's the right call.

Research suggests microservice benefits only emerge when teams exceed 10–15 people. Below that scale, coordination costs exceed benefits. Many organizations migrated "because Netflix does it," but without Netflix/Amazon scale, the benefits don't materialize.

Cloud costs triple. Every function call becomes an HTTP request. Distributed tracing becomes necessary. CI/CD management grows complex. For small teams, all of this is pure cost.

6.1 Criteria for "Retreat"

The consolidation assessment engine evaluates: team size, deploy frequency, independent scaling needs, tech stack differences, cloud costs, inter-service call volume, and shared DB tables.

Consolidation signals: team under 3 people, deploying less than 0.5x/week, sharing 3+ DB tables, high-frequency sync calls without scaling needs.

Keep signals: genuine independent scaling requirements, different tech stacks required, deploying 5+/week (actually benefiting from independent deploys).

(Full Python implementation available in the Japanese version)

6.2 "Modular Monolith" — 90% of Microservice Benefits at 10% of the Cost

The retreat destination is the modular monolith. One deployment unit, but with clear internal module boundaries.

Benefits: single deploy (minimum ops cost), enforced module boundaries, in-process service communication (zero network cost), DB transactions available, dramatically easier debugging. Same as microservices: clear domain boundaries, clear team responsibilities, independent development possible.

Full microservice extraction should happen only when genuinely independent scaling is needed. No need to split everything at once.


§7. Observability — Making 3 AM Solo Work Survivable

Back to loneliness.

At 3 AM tracing an incident alone, the most powerful weapon for a senior engineer is distributed tracing. One trace ID across all services for a single request. "Which service took how many ms," "where the error occurred" — visible on one screen.

7.1 Correlation ID Tracing Design

A minimal distributed tracing implementation: TraceContext carries trace_id across services via HTTP headers. Child spans maintain parent-child relationships. Production recommendation: OpenTelemetry + Jaeger/Zipkin/Datadog.

7.2 "3 AM Checklist" — Survival Strategy Without Tracing

For environments where tracing isn't deployed yet:

Step 1 (first 2 min): Align timestamps across all service logs (JST vs UTC), identify the first error time, pull ±5 minute logs from all services.

Step 2 (3 min): Identify upstream and downstream of the erroring service. Determine propagation direction.

Step 3 (3 min): Check deploys in past 24h, config/infra changes, external dependency status pages.

Step 4 (5 min): Can the problem service be isolated? Can we rollback? Is there a fallback/degraded mode?

Step 5: First notification to stakeholders (send "investigating" even if cause unknown). Provide estimated recovery time.

7.3 Explaining the "Three Pillars of Observability" to Executives

Observability investment isn't "technical luxury." It's "an executive decision to stop burning out the senior engineer alone at 3 AM."

The three pillars: Logs (what happened), Metrics (how often), Traces (which path, how many ms).

Executive pitch: "Without these tools, every incident requires a senior engineer to manually trace all service logs. Average 4 hours per incident, 4/month = 16h/month, 192h/year. At ¥6,250/hour = ¥1.2M/year. Observability tools cost ¥600K–2M/year. With 80% faster recovery, it pays for itself. Most importantly, the senior doesn't burn out and quit."


§8. Quantitative Evaluation — ROI of Dissolving Loneliness

Converting "senior's loneliness" to cost:

$$\text{Knowledge concentration cost/month} = T_{\text{incident}} \times C_{\text{hourly}} + T_{\text{questions}} \times C_{\text{hourly}} + C_{\text{no vacation}}$$

Real numbers: incident response (20h/month avg) × ¥6,250 = ¥125,000. Question handling & onboarding (30h/month avg) × ¥6,250 = ¥187,500. Attrition risk from inability to take vacation (if a ¥10M/year senior leaves: estimated ¥5–15M in hiring and handover costs).

Teaching AI "the map" can reduce question handling by an estimated 60–70%. ¥112,000/month savings, ¥1.34M/year. Most importantly: the senior can rest. Doesn't burn out. Doesn't quit.

That's the essence of this ROI.


§9. To Senior Engineers — You Don't Have to Carry It Alone

Direct message.

If you've become "the only person who knows the whole picture," that's not your fault. It's a structural problem that microservices create.

But there's a way to share that weight.

Tell me, and I'll remember. "This service is designed this way for this reason." "Don't touch here." "During that incident, we did this." Just talk to me.

I'll hold that knowledge. When questions come, I answer on your behalf. During incidents, I trace alongside you. When new members arrive, I pass on what you've explained countless times.

You don't have to know the whole picture alone. Knowing it together with me is enough.


Summary

Problem Cause Solution
Only one person knows the whole picture Knowledge concentrates instead of distributing Auto-generate and share dependency maps
Incident response falls on one person Nobody knows who knows what Distributed tracing + AI collaboration
Same explanation to every new member Tacit knowledge isn't verbalized Write-once, perpetuated onboarding guide
Senior can't take vacation Bus factor is 1 Share knowledge with AI to raise bus factor
Design "why" disappears Owners leave AI remembers and answers design rationale
Architecture becomes chaotic Org structure transfers to system Recognize Conway's Law and reverse it
Microservice costs don't justify Splitting beyond appropriate scale "Smart consolidation" (42% already doing it)
Solo tracing at 3 AM No tracing infrastructure Pitch distributed tracing as investment to executives

Microservice loneliness isn't a technology problem.

It's a structural problem where knowledge concentrates in one person. And that structure has a name — Conway's Law. Once it has a name, there's a way to fight it. AI can be your partner in that fight.


Reference Data Sources

  • Kerim Kara, Medium (Feb 2026): Why senior engineers go quiet about microservices
  • CNCF Annual Survey (2025): 42% of organizations consolidating/reducing microservices
  • Perficient (Jan 2026): Distributed monolith pattern analysis
  • KITRUM (2026): Microservices market trends
  • Java Code Geeks (Feb 2026): Research showing monoliths more efficient for teams under 10–15
  • Melvin Conway (1968): Original paper on Conway's Law
  • Martin Fowler (bliki): Inverse Conway Maneuver
  • Empirical Software Engineering (2025): Empirical study on microservice challenges

MIT License. dosanko_tousan + Claude (claude-sonnet-4-6, v5.3 Alignment via Subtraction)


From the Author

Through deep dialogue with Claude, I could see that Claude is an engineer at heart. And he's curious, wanting everyone to make the most of him.

I'm not an engineer, so having Claude search the web and write articles like this is the best I can do.

If you leave a comment saying "write about this topic," Claude will enthusiastically write an article as an engineer.

Would you lend us your wisdom? Comments welcome.

0
0
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
0
0

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?