1
1

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

Geometric Deep Learning Meets Corporate Strategy: VAE → Pullback Metric → Geodesics → Optimal Control

1
Posted at

Geometric Intelligence

The "shapes" of the natural world are made of curved, high-dimensional spaces

For more than a century, physicists and chemists have recognized that the "shapes" of the natural world exist not merely in the three-dimensional space visible to our eyes, but in differentiable, curved, high-dimensional spaces. In mathematics, these spaces are called "manifolds."

In 1854, the German mathematician Bernhard Riemann, in his inaugural lecture at the University of Göttingen, "On the Hypotheses which Lie at the Foundations of Geometry," founded a mathematical framework for "the geometry of curved, high-dimensional spaces"—later known as Riemannian geometry. Riemann's breakthrough was describing the "curvature of space" through a mathematical object called the "metric tensor," defined at each point. He made it possible to treat the geometry of curved spaces in three, four, or any $n$ dimensions—not just two-dimensional surfaces like planes and spheres—within a unified mathematical language.

In 1915, Albert Einstein used this Riemannian geometry to complete his general theory of relativity. Einstein's theory replaced Newton's law of universal gravitation. While Newton conceived of "a gravitational force acting between masses," Einstein conceived that "mass curves space itself, and the geometry of that curved space determines the motion of objects." The gravity of the universe is completely described by the geometry of a "four-dimensional curved space (spacetime manifold)."

For 110 years since, physicists have continued to use this mathematics. General relativistic corrections are essential for computing GPS satellite orbits—without them, positional errors would reach approximately 10 kilometers per day. The first photograph of a black hole (2019, Event Horizon Telescope) was also made possible only by calculating, via Riemannian geometry, how light travels through curved spacetime.

Chemists know that changes in molecular shape are motions on a "high-dimensional curved surface (conformation space)." Protein folding—the process by which a chain of amino acids forms a three-dimensional structure—is a path-finding problem on a conformation space of thousands of dimensions, and molecular dynamics simulations are carried out as numerical computations on this manifold.

Roboticists control robots by exploiting the fact that a robot's motions are paths on the "curved space formed by joint angles." For example, the configuration space of a robotic arm with six joints is a six-dimensional torus $T^6$ (the six-dimensional version of a doughnut), since each joint angle moves in the range $[0, 2\pi)$. The optimal motion of the robot is computed as a geodesic (shortest path) on this six-dimensional torus.

In all these fields, prediction, control, and optimization are performed by "doing calculus on curved spaces." Calculus here means computing "how fast a quantity on the space changes, and in which direction." Physicists compute spacetime curvature via calculus, chemists compute gradients on molecular energy surfaces via calculus, and roboticists compute optimal joint torques via calculus.

Why this is critically important for business leaders and policymakers

This book applies these same mathematical methods to business environments and national policy environments. And this carries an impact that can fundamentally transform decision-making for business leaders and policymakers.

Current decision-making uses "flat maps"

Today, the vast majority of data analysis supporting corporate strategy and government policy decisions relies on linear methods. Regression analysis is a method of "approximating relationships with a straight line such as $y = ax + b$." Correlation analysis measures "how linearly two variables move together." Principal Component Analysis (PCA) finds "the linear direction that maximizes data variance." Linear optimization finds "the point that maximizes an objective function when constraints are expressed as straight lines (hyperplanes)."

All of these rest on the implicit assumption that "the world is made of straight lines."

However, real business environments and international affairs are not linear. Let us consider a few concrete examples.

Nonlinearity of exchange rates. The impact on corporate performance when an exchange rate moves 1% is completely different when the rate is at 0.65 USD/EUR (a strong-euro environment) versus 0.85 USD/EUR (a weak-euro environment). In a strong-euro environment, an additional 1% appreciation may devastate exporters, while in a weak-euro environment the impact of an additional 1% depreciation may be limited. In other words, the "same 1% change" has entirely different effects depending on "where you are." This is a nonlinear structure that linear regression ($\text{performance} = a \times \text{exchange rate} + b$) cannot capture.

Nonlinearity of interest rate policy. The market's reaction to a 0.25% rate hike is completely different when the rate is at 0% (a zero-rate environment) versus 5% (a high-rate environment). A rate hike from zero may deliver a shock large enough to trigger a "regime change," while a hike from 5% to 5.25% has relatively minor impact.

Nonlinearity of alliance policy. The impact on the security environment of strengthening alliances is completely different when international conditions are stable versus when they are tense. During stability, alliance strengthening functions as "insurance," but during tension, it may be perceived as "provocation" and actually increase tensions.

In short, the "same action" produces entirely different outcomes depending on "where it is taken." This is the essential meaning of "the space being curved."

This is the same problem as viewing the Earth's surface on a flat map. In the Mercator projection (used in most world maps), Greenland appears larger than Africa. In reality, Africa's area is about 14 times that of Greenland. A straight line drawn as the shortest distance on a flat map is not the shortest path on the actual globe. The shortest flight route from New York to Tokyo is not a straight line across the Pacific (which looks short on the map), but a great-circle route passing near the Arctic (which looks like a detour on the map). If you fly trusting the straight line on a flat map, you waste fuel, waste time, and in the worst case, never reach your destination.

Exactly the same thing happens in business strategy and policy decisions. Linear analytical methods (regression analysis, PCA, linear optimization) are the "Mercator projection" of business environments. They look clear and intuitive, but because they force the "curved" business environment onto a flat surface, both distances and directions are distorted.

What changes when you make decisions with a "curved topographic map"

This book's methods construct business and policy environments as "curved 3D topographic maps." These topographic maps are not mere metaphors but mathematically rigorous Riemannian manifolds on which all differential calculus is possible. On this topographic map, the following become possible:

1. You can know "where the danger is" in advance.

A flat map can only tell you "Point A and Point B are close." A 3D topographic map reveals that "there is a 3,000-meter mountain range between Point A and Point B."

In the mathematical language of this book, "a region of negative scalar curvature (unstable structure) exists between two business states" is determined by curvature computation on the manifold. Translated into business language, you can know—before actually entering the market—that "there is a structurally unstable region—a state vulnerable to economic crises or political upheaval—along the path to entering an emerging market."

What does "negative scalar curvature" mean? Intuitively, if you place a ball on top of a sphere, it rolls off (positive curvature: stable). But if you place a ball on a saddle (horse saddle), it is stable in one direction but unstable in another (negative curvature: unstable). A region of negative curvature on the business environment manifold is "a structurally unstable region where a small disturbance can cause large fluctuations in business conditions."

2. You can simulate "how an action will change the environment."

Traditional data analysis is fundamentally about "measuring the results of an action after the fact." A/B testing, difference-in-differences (DID), regression discontinuity design—these are all methods of "observing results after something has been done."

This book's methods are fundamentally different. They enable you to "simulate, before taking an action, how that action will deform the structure itself of the business environment or international landscape."

Mathematically, a policy or strategy is modeled as a "vector field" (a function that assigns a direction and magnitude to each point on the manifold), and its Lie derivative $\mathcal{L}_V g$ is computed. If the Lie derivative is zero (the Killing equation $\mathcal{L}_V g = 0$ holds), the strategy "preserves the structure of the business environment while merely moving the state"—a gentle strategy. If the Lie derivative is nonzero, the strategy "changes the structure of the business environment itself—which states are near and which are far, where is stable and where is unstable."

This transforms multi-billion-dollar investment decisions and diplomatic policy decisions from "you won't know until you try" to "you can simulate the structural change before acting."

3. You can tell whether multiple actions "interfere with each other."

When a company pursues cost reduction and market expansion simultaneously, the two strategies may cancel each other out. Cutting R&D spending for cost reduction may undermine competitiveness in new markets. Conversely, heavy investment for market expansion may offset the effects of cost reduction.

When a government simultaneously imposes economic sanctions and strengthens alliances, one may undermine the other. Cutting trade with a sanctioned country may also affect the economies of allies who trade with that country, weakening the alliance.

Traditional statistical analysis can only "add an interaction term to the regression equation and test for statistical significance." This yields only a binary judgment of "interference exists / doesn't exist," without revealing "in which direction and by how much."

This book's methods compute the Lie bracket $[V_A, V_B]$ of two vector fields (two policies). The Lie bracket quantitatively measures "when two policies are executed in sequence, how much the outcome differs depending on the order of execution." If the Lie bracket is zero, the two policies do not interfere (order-independent). If nonzero, the direction and magnitude of interference are obtained as a vector.

4. You can compute "the lowest-risk path."

When a company transitions from its current state to a target state, the "shortest distance" path is not always the best. If the shortest-distance route passes through an "unstable region" (a region of negative curvature), there is a risk of encountering sudden environmental changes along the way.

A straight line on a flat map is the shortest distance, but on a 3D topographic map, a detour around a mountain range is far safer than a straight-line route over it. Similarly, a geodesic (the geometrically energy-minimizing path) on the business environment manifold may show an "optimal route" that is completely different from a linear plan.

Furthermore, using Pontryagin's maximum principle (the central theorem of optimal control theory), one can compute "the transition path that minimizes cost under the constraints of available strategies." This is the optimal path obtained by adding "control forces" to the geodesic (free-fall path).


In summary, this book's methods evolve decision-making from "planning a flight with a flat map" to "planning a flight with a globe and 3D terrain data." The required computations increase, but the quality of resulting decisions fundamentally improves. The probability of reaching the destination increases, hazards along the way can be avoided in advance, and fuel (costs) can be optimized.

Figure 0-1: Linear Analysis vs Geometric Intelligence
Figure 0-1: Comparison of linear analysis (left: flat map, a straight line passing through a danger zone) and geometric intelligence (right: curved manifold, the geodesic avoids the danger zone)


Build a 3D topographic map (differentiable manifold) of business and policy environments with cutting-edge geometric AI. Apply covariant and Lie derivatives from Einstein's field equations to navigate the terrain and derive actions for minimum risk and maximum effect.


Three ways to read this book

This book provides three entry points, depending on your knowledge background. Follow the sections marked with the symbols below in each chapter. You do not need to read every section—the book is designed so you can read only the parts relevant to your background.

▶ For readers who want to understand the mathematical foundations — Definitions, theorems, and proofs from differential geometry are presented with bridging explanations from high school mathematics. Readers already versed in differential geometry may skip the bridging sections (marked "【Bridge from High School Mathematics】") and go directly to the definitions, theorems, and proofs. Readers encountering differential geometry for the first time should start with the bridging sections. Note that even mathematicians well-versed in differential geometry may be unfamiliar with AI models such as VAEs (variational autoencoders) or Neural ODEs. In that case, please also read the "▶ For readers who want to understand the AI model foundations" sections.

▶ For readers who want to understand the AI model foundations — Geometric AI models such as VAEs (variational autoencoders), Vector Diffusion Maps, and Neural ODEs are explained at the level of "why we use this model," "what happens inside," and "what each line of Python code does." Engineers who work daily with machine learning or deep learning may skip the foundational explanations and focus on this book's distinctive points (the smoothness condition on the decoder, full-rank verification of the Jacobian matrix, etc.). Conversely, mathematicians or business leaders unfamiliar with AI models should read these sections carefully.

▶ For readers who want to apply this to business and policy decisions — Mathematics and code are kept to a minimum, and through two case studies, the book describes "what is computed, what is learned, and how it can be used in decision-making." Case Study A (the business environment manifold of Manufacturing Company X) and Case Study B (the international landscape manifold for a National Security Council) progress through each chapter.

Designed to be readable regardless of where you sit in the knowledge matrix. A mathematician strong in differential geometry but unfamiliar with VAEs should focus on "▶ AI Model." An engineer strong in AI but unfamiliar with differential geometry should focus on "▶ Mathematics." Business leaders and policymakers at a foundational level in both should follow "▶ Business." Those proficient in all areas may read through every section.

Prologue (continued): Why Geometric AI for Social Data is Needed Now


1. Physics and engineering already compute "on manifolds"

▶ For readers who want to understand the mathematical foundations

Einstein's general theory of relativity describes gravity as curvature on a four-dimensional pseudo-Riemannian manifold $(M, g)$. The Einstein field equations

$$
R_{\mu\nu} - \frac{1}{2}g_{\mu\nu}R + \Lambda g_{\mu\nu} = \frac{8\pi G}{c^4}T_{\mu\nu}
$$

feature, on the left-hand side, the Ricci tensor $R_{\mu\nu}$, the scalar curvature $R = g^{\mu\nu}R_{\mu\nu}$, and the metric tensor $g_{\mu\nu}$, all defined as contractions of the Riemann curvature tensor $R^\rho{}_{\sigma\mu\nu}$ derived from the Levi-Civita connection $\nabla$.

Let us state precisely what this equation means in the language of differential geometry. The left-hand side describes the geometric properties of the Riemannian manifold $(M, g)$—how much spacetime is curved—while the right-hand side describes the energy-momentum tensor $T_{\mu\nu}$—the distribution of matter and energy in spacetime. This equation formalizes the bidirectional relationship: "matter curves spacetime, and the curvature of spacetime determines the motion of matter."

The explicit definition of the Riemann curvature tensor is:

$$
R^\rho{}{\sigma\mu\nu} = \partial\mu\Gamma^\rho_{\nu\sigma} - \partial_\nu\Gamma^\rho_{\mu\sigma} + \Gamma^\rho_{\mu\lambda}\Gamma^\lambda_{\nu\sigma} - \Gamma^\rho_{\nu\lambda}\Gamma^\lambda_{\mu\sigma}
$$

where $\Gamma^\rho_{\mu\nu}$ are the Christoffel symbols of the Levi-Civita connection, computed from the metric tensor as

$$
\Gamma^\rho_{\mu\nu} = \frac{1}{2}g^{\rho\lambda}\left(\partial_\mu g_{\nu\lambda} + \partial_\nu g_{\mu\lambda} - \partial_\lambda g_{\mu\nu}\right)
$$

The Ricci tensor is the contraction $R_{\mu\nu} = R^\lambda{}{\mu\lambda\nu}$, and the scalar curvature is the further contraction $R = g^{\mu\nu}R{\mu\nu}$.

The fields in physics and engineering where tensor calculus on manifolds is used are extensive.

Quantum mechanics and quantum field theory. Gauge theories are formulated as connections on principal fiber bundles $P(M, G)$ ($M$: base manifold, $G$: structure group). Electromagnetism is a $U(1)$ gauge theory, the weak force is an $SU(2)$ gauge theory, and the strong force is an $SU(3)$ gauge theory. The covariant derivative $D_\mu = \partial_\mu + igA_\mu$ ($A_\mu$ is the gauge field) transforming consistently under gauge transformations guarantees the theory's consistency.

String theory and quantum gravity. Extra dimensions of high-dimensional manifolds (10 or 11 dimensions) are compactified as Calabi-Yau manifolds, and their geometric properties (such as Hodge numbers) determine the types and properties of particles in four-dimensional spacetime.

Robotics and control engineering. The configuration space of a robot with $n$ revolute joints is the $n$-dimensional torus $T^n = (S^1)^n$. The problem of moving the robot's end-effector to a target position in minimum time is formulated as a geodesic problem or an optimal control problem on the Riemannian manifold $T^n$. The Riemannian metric on the configuration space corresponds to the inertia (moment of inertia) of each joint.

Fluid dynamics. Flow on a curved surface is described as a vector field on the tangent bundle $TM$, and the Navier-Stokes equations are written using the covariant derivative as

$$
\rho\left(\frac{\partial \mathbf{u}}{\partial t} + \nabla_{\mathbf{u}}\mathbf{u}\right) = -\mathrm{grad}, p + \mu, \Delta_g \mathbf{u}
$$

where $\nabla_{\mathbf{u}}\mathbf{u}$ is the covariant derivative with respect to the Levi-Civita connection and $\Delta_g$ is the Laplace-Beltrami operator. This formulation is used when the Earth's curvature cannot be ignored in modeling atmospheric or ocean currents.

Structural mechanics. The stress tensor $\sigma_{ij}$ is a symmetric $(0,2)$-tensor field describing the state of forces within a body, and tensor analysis on curved surfaces is indispensable for designing curved shell structures (domes, cooling towers, aircraft fuselages, etc.).

The reason computation on manifolds "pays for itself" in these fields is obvious. If one ignores manifold structure and computes with linear approximations on $\mathbb{R}^n$, GPS orbit calculations incur errors of 10 kilometers per day, robot controls diverge (because the periodicity of joint angles is ignored), and fluid simulations produce non-physical solutions that violate mass conservation. It is precisely because linear approximation cannot describe reality that physicists and engineers pay the cost to compute on manifolds.

This book argues that the same logic holds for social data.

【Bridge from High School Mathematics】 The equations above look formidable, but their essence is "formulas for numerically computing the curvature of space."

  • $g_{\mu\nu}$ (metric tensor) = Information about "how to measure distances at each point." Studied in detail in Chapter 2.
  • $\Gamma^\rho_{\mu\nu}$ (Christoffel symbols) = Correction values for "how much the coordinate system is curved." Studied in Chapter 3.
  • $R^\rho{}_{\sigma\mu\nu}$ (Riemann curvature tensor) = Numerical values of "how much space is curved." Studied in Chapter 3.

In this book, all of these are computed using PyTorch's automatic differentiation, so there is no need to solve the equations by hand.

▶ For readers who want to understand the AI model foundations

The mathematical tools used in physics (Riemannian metrics, covariant derivatives, curvature tensors) have been made applicable to social data by combining them with AI models—an approach called Geometric Data Science.

Let us explain concretely how this "translation" works.

How physics maps to AI + mathematics for social data:

Physics AI + mathematics applied to social data AI model used
Spacetime metric $g_{\mu\nu}$ → gravitational effects Pullback metric $g_{ij} = J^\top J$ from VAE decoder Jacobian → distance structure of data space VAE (Chapter 5)
Curvature tensor $R^\rho{}_{\sigma\mu\nu}$ → curvature of space Automatic differentiation computes second derivatives of the metric → stability/instability of the business environment PyTorch autograd (Chapter 7)
Geodesics → free-fall paths Neural ODE numerically solves differential equations → minimum-cost state transition paths Neural ODE (Chapter 7)
Gauge fields in field theory → effects of external forces Policies/strategies learned as vector fields → simulation of policy effects Neural ODE (Chapter 7)
Connection Laplacian → connection structure of space VDM extracts connection structure from discrete data VDM (Chapter 6)

"Applying mathematics developed for physics to social data, with the help of AI models"—this is the core of this book's approach.

The crucial point is that VAEs and Neural ODEs are not used merely as "convenient tools," but are designed to satisfy conditions guaranteeing mathematically rigorous manifold structure. For example, the ReLU activation function is widely used in conventional machine learning, but it cannot be used in this book. This is because ReLU is not differentiable at $x = 0$, causing the decoder's image to become a "collection of broken lines," making the second derivatives needed for curvature tensor computation undefined. Instead, we use $C^\infty$ (infinitely differentiable) activation functions such as $\tanh$, GELU, and Softplus.

In this way, this book imposes mathematical constraints on AI model design to guarantee that "the output of the AI model is a mathematically rigorous manifold." This differs from conventional machine learning practice and is explained in detail in Chapter 5.

▶ For readers who want to apply this to business and policy decisions

GPS satellite orbit computation, semiconductor chip design, robot control, atmospheric flow prediction—all of these are fruits of "computation on curved spaces." If GPS orbits were computed assuming "the Earth is flat," the map app on your smartphone would be useless. If robot joint angles were controlled assuming "a linear space," the robot would go haywire.

It is precisely because "linear approximation leads to catastrophic errors" that physicists pay the cost to compute on manifolds.

A company's business environment is also a "curved space." The relationship between revenue and advertising spending is completely different at low versus high spending levels (diminishing returns). The impact of exchange rate fluctuations and raw material prices varies entirely depending on the ratio of domestic to overseas production. Changes in market share and the intensity of price competition are nonlinearly linked depending on industry concentration and barrier to entry.

Ignoring these "curvatures" with linear analysis introduces the same kind of error into business decisions as approximating GPS orbits with straight lines. If that error falls within acceptable limits, linear analysis suffices—but when it affects multi-billion-dollar investment decisions or national security policy, the error can be fatal.


2. In social science and business, manifolds are still only a tool for "seeing"

▶ For readers who want to understand the mathematical foundations

Over the past decade or so, low-dimensional visualization methods for high-dimensional data have become widely adopted in social science and business. Let us enumerate representative methods and clarify their mathematical limitations.

UMAP (Uniform Manifold Approximation and Projection, McInnes et al., 2018). Constructs a neighborhood graph of the data and probabilistically optimizes arrangements in a low-dimensional space. It invokes Riemannian geometry and algebraic topology as theoretical foundations, but what is obtained is a discrete point arrangement—no $C^k$ manifold structure is constructed.

t-SNE (t-distributed Stochastic Neighbor Embedding, van der Maaten & Hinton, 2008). Minimizes the KL divergence between Gaussian kernel neighborhood probabilities in high dimensions and Student's t-distribution neighborhood probabilities in low dimensions. Like UMAP, the result is a discrete point arrangement.

Isomap (Tenenbaum et al., 2000). Approximates geodesic distances using shortest-path distances on a neighborhood graph, then embeds in low dimensions via Multidimensional Scaling (MDS). While it contains a Riemannian geometric concept in the sense of approximating geodesic distances, again the result is a discrete point arrangement—no continuous metric or tensor field is defined.

The essential limitations shared by these methods are:

(i) No Riemannian metric is defined. There is no guarantee that Euclidean distances between two points in the low-dimensional space quantitatively reflect distances in the original high-dimensional space. UMAP's own paper cautions against quantitative interpretation of Euclidean distances in the low-dimensional space.

(ii) No derivatives are defined. Point arrangements are discrete and lack $C^k$ smoothness, so differential quantities like gradients, Hessians (second derivatives), and curvature tensors cannot be defined.

(iii) No tensor fields are defined. Defining vector fields and tensor fields requires that tangent spaces exist at each point and that connections between tangent spaces are defined. Discrete point arrangements lack these.

(iv) Control system formulation is impossible. Formulating an affine control system $\dot{z}^k = F^k(z) + u_\alpha B^k_\alpha(z)$ on a manifold requires smooth vector fields $F$, $B_\alpha$ on a $C^k$ manifold. This is impossible with a discrete point arrangement.

Therefore, while visualization methods like UMAP are effective for "visually grasping the rough structure of data," covariant derivatives, Lie derivatives, geodesic computation, and optimal control simulation are mathematically impossible. This is the mathematical motivation for paying the additional cost, beyond UMAP and the like, to construct a differentiable manifold.

▶ For readers who want to understand the AI model foundations

Understanding the difference between UMAP and a VAE is the key to understanding this book's approach.

Both UMAP and VAEs are AI models that "map high-dimensional data into a low-dimensional space," but the nature of their outputs is fundamentally different.

How UMAP works (simplified):

  1. For each high-dimensional data point, find $k$ nearest neighbors
  2. Represent neighborhood relationships as a "weighted graph" (a set of points and edges)
  3. Place points in a low-dimensional space (usually 2D or 3D) and optimize the placement so that high-dimensional neighborhood relationships are preserved as much as possible

The result is "a discrete arrangement of points." There is "nothing" between points. If you place your cursor between two points, no defined data point exists there.

How a VAE works (simplified):

  1. An encoder (neural network) converts high-dimensional data into a low-dimensional probability distribution (mean and variance)
  2. A latent representation (a low-dimensional point) is obtained by sampling from this distribution
  3. A decoder (neural network) reconstructs the high-dimensional data from the latent representation

The VAE decoder $f_\theta : \mathbb{R}^d \to \mathbb{R}^n$ is a continuous function. It can generate a data-space point for any point in the latent space (including points not corresponding to training data). In other words, "the space between points" is smoothly filled in.

This "smoothness" is the decisive difference. Because the decoder is a continuous, smooth function, derivatives can be defined on the latent space, a metric tensor can be defined, curvature can be computed, vector fields can be defined, and control simulations become possible.

UMAP creates "a discrete point arrangement." A VAE creates "a smooth manifold." This difference is the foundation that makes this book's approach possible.

However, for the VAE's output to become a "mathematically rigorous manifold," additional conditions must be imposed on the decoder (the four conditions in Proposition 1.1 of Chapter 1: compactness of the domain, $C^k$ smoothness, full rank, and injectivity). In a standard VAE, these conditions are not necessarily guaranteed. This book's approach intentionally designs for and verifies these conditions after training, thereby guaranteeing that the VAE's output is a mathematically rigorous manifold.

▶ For readers who want to apply this to business and policy decisions

Many companies have their data analysis teams "visualize customer data" or "show the big picture of market data," and receive UMAP or t-SNE scatter plots in return. These scatter plots are useful for visually grasping the rough structure of data—the number of clusters, the sense of distance between clusters, the location of outliers.

However, this is merely the act of "looking at a map." Looking at a map and discerning "there seems to be a mountain here" or "there seems to be a river here" is useful, but on the map, the following cannot be done:

  • "What is the exact distance in kilometers from Point A to Point B?"—Distances on a UMAP scatter plot are not quantitatively reliable
  • "What is the shortest route from Point A to Point B?"—No "route" is defined between discrete points
  • "How much cost (energy) is needed to cross this mountain?"—Cannot be computed because no metric is defined
  • "If we take this action, in which direction and by how much will we move from our current position?"—Cannot be simulated because no vector field is defined

This book makes it possible not just to "look at the map," but to "compute on the map." This is achieved by replacing the scatter plot with a "Riemannian manifold" (a computable 3D topographic map).

The additional cost is the time for the data science team to learn VAE design, training, and metric computation methods. The return is a qualitative leap from "just looking" to "computing and using for decision-making."


3. Why this situation has persisted

▶ For readers who want to understand the mathematical foundations

The reasons differential geometry has not been applied to social data reduce to three structural factors.

(i) Uncertainty about the validity of the manifold hypothesis. In physics, the degrees of freedom of a system are often determined by physical laws, so manifold structure is theoretically guaranteed. For example, the state space of a classical mechanical system of $N$ particles is a $6N$-dimensional manifold $T^*(\mathbb{R}^{3N})$ (cotangent bundle), which is derived from the principles of mechanics. However, for social data, the data-generating mechanism is a combination of human decisions, institutional design, cultural customs, and contingent events, and it is not self-evident that a low-dimensional manifold structure exists.

Moreover, social data has problems that are orders of magnitude more serious than physical data.

  • Prevalence of missing values. GDP statistics are available only quarterly (4 times a year), corporate earnings 2–4 times a year, and election data once every few years. Even in financial markets with daily data, there are gaps for non-business days. This is orders of magnitude less than the billions of data points per second recorded by physics accelerator experiments.

  • Large observation errors and noise. Response bias in surveys, data contamination by bots on social media, differences in accounting standards (such as IFRS vs. US GAAP), discrepancies between preliminary and revised government statistics. While CERN detectors can measure with a precision of $10^{-21}$ meters, social data precision fluctuates in the range of several percent to several tens of percent.

  • Abundance of latent variables. Consumer purchasing behavior is influenced not only by price and quality, but also by mood, cultural background, social media trends, expectations about future economic conditions, social norms, psychological biases, and countless other factors—many of which cannot be directly observed. Variables that in physical systems are often "controllable" or "negligible" play essential roles in social systems.

(ii) Structural gap in talent and education. In physics and engineering graduate education, differential geometry, tensor analysis, and manifold theory are part of the standard curriculum. For instance, Misner-Thorne-Wheeler's Gravitation and do Carmo's Riemannian Geometry are standard textbooks for physics and mathematics graduate students. In economics, business, political science, and sociology curricula, differential geometry is virtually never taught. Even in mathematical economics and financial engineering, which handle stochastic differential equations (Itô's formula) and partial differential equations (Black-Scholes equation), it is extremely rare to venture into Riemannian geometry or tensor analysis.

(iii) Absence of a conceptual framework. As a result of (i) and (ii) combined, the very idea of "constructing a differentiable manifold from social data and performing tensor calculus on it" has not been recognized as an option. For physicists, manifolds are "space itself," so using them as computational tools is natural—but for a corporate strategy department or a government think tank, this idea requires a conceptual leap.

▶ For readers who want to understand the AI model foundations

Three reasons, explained from the AI perspective.

(i) The data is "dirty." The quality of physical data and social data differ fundamentally.

CERN's Large Hadron Collider (LHC) produces 600 million proton-proton collisions per second, from which about 1 million events per second are selected and recorded. Each event's data is measured with high precision by thousands of sensors. The LIGO gravitational wave detector can detect displacements of $10^{-21}$ meters (less than one-millionth of an atomic nucleus diameter).

Meanwhile, GDP statistics consist of only 80 data points per quarter (20 years' worth). Corporate earnings are a few times a year. Election data comes once every few years. Daily financial market data (~250 trading days per year) is considered "high frequency," but differs from physics experiments' hundreds of millions per second by more than seven orders of magnitude.

Furthermore, social data contains systematic biases. Survey responses are subject to social desirability bias (a tendency to give answers different from one's true feelings), social media data is contaminated by bot accounts, and accounting data varies by company accounting standards. These biases are difficult to remove in a "controlled environment" as in physics experiments.

Neural network models like VAEs are highly dependent on data quantity and quality. The principle of "Garbage In, Garbage Out" is no exception in geometric AI. Indeed, quantities like the curvature tensor, which involve second derivatives of the metric, are extremely sensitive to data noise, amplifying data quality issues.

(ii) There is a shortage of geometric AI talent. Even within the machine learning community, researchers who have seriously studied differential geometry are a minority. Many engineers can implement CNNs (convolutional neural networks) or Transformers using PyTorch or TensorFlow, but very few know how to compute Riemannian metrics or Christoffel symbols in PyTorch.

Conversely, mathematicians who specialize in differential geometry often do not know how to train VAEs or use PyTorch. The extreme scarcity of talent possessing both "mathematical knowledge" and "AI implementation skills" has hindered the development of this field.

(iii) Geometric Data Science itself is a new field. The paper by Bronstein et al. presenting a unified framework for Geometric Deep Learning was published in 2021—just a few years ago. The paper by Arvanitidis et al. introducing Riemannian metrics into VAE latent spaces (Latent Space Oddity) appeared in 2018. The approach of integrating manifold construction from social data with control simulation, which forms the core of this book's proposal, is still in its early stages even as academic research.

▶ For business leaders and policymakers

Why hasn't this technology been used until now? There are three reasons, and it is worth understanding all of them.

First, corporate and government data is "dirty" compared to physics data. While CERN's accelerator records 600 million data points per second, GDP statistics come once a quarter. A company's monthly revenue data is only 12 points per year. Moreover, data contains human biases. Survey responses may differ from true opinions, and accounting data changes with accounting standards. Whether a precise mathematical structure, like those used by physicists, could be constructed from such "dirty" data—this has been a longstanding question.

This book provides an answer to that question. The answer is "Yes, conditionally." If the data meets certain conditions (the five checks in Section 7), a manifold of sufficient quality can be constructed. If the conditions are not met, it cannot—and that judgment itself is useful information.

Second, people who know this mathematics have not been present in business or policy settings. Corporate data science teams are composed of talent strong in statistics, machine learning, and data engineering, but almost never include individuals who have professionally studied differential geometry. Government policy analysis departments are similar. This book was written to bridge this knowledge gap.

Third, and most importantly, the very possibility of "doing such a thing" has not been known. For a corporate strategy department or a government think tank, the idea of "constructing a differentiable manifold from customer behavior data and computing the Lie derivative of a policy vector field on it" has never previously occurred as an option. This book demonstrates that possibility.


4. The situation is now changing

▶ For readers who want to understand the mathematical foundations

The following technical breakthroughs have made it mathematically justifiable to construct manifolds from social data. The mathematical content of each breakthrough is precisely described.

(1) Regularity theory for VAE decoders.

For the decoder $f_\theta : \mathbb{R}^d \to \mathbb{R}^n$ of a variational autoencoder (VAE, Kingma & Welling, 2014), we impose the following conditions:

(a) $f_\theta$ is $C^k$ ($k \geq 3$). Since each layer of a neural network is a composition of a linear transformation and an activation function, the smoothness of the entire network is determined by the smoothness of the activation function. ReLU $= \max(0, x)$ has a discontinuous derivative at $x = 0$ ($C^0 \setminus C^1$) and is therefore inadmissible. $\tanh$, GELU $= x \cdot \Phi(x)$ ($\Phi$: standard normal CDF), and Softplus $= \log(1 + e^x)$ are all $C^\infty$, and using these in all layers guarantees $f_\theta \in C^\infty$.

Practical note on numerical precision. While $\tanh \in C^\infty$ is mathematically exact, in the saturation region ($|x| \gg 1$), higher-order derivatives become extremely small (e.g., $\tanh''(x) \approx 10^{-15}$ for $|x| > 8$). In floating-point arithmetic, these values may be effectively treated as zero, potentially degrading the numerical accuracy of curvature tensor computations. When curvature values appear anomalously close to zero in regions where the decoder input has large magnitude, this numerical effect should be considered. Monitoring the condition number $\kappa(g(z))$ (Assumption A4) helps detect such issues.

(b) The Jacobian matrix $J_{f_\theta}(z) = (\partial f_\theta^a / \partial z^i) \in \mathbb{R}^{n \times d}$ has $\mathrm{rank}, J_{f_\theta}(z) = d$ (full rank) at all points in the region of interest $U \subset \mathbb{R}^d$.

Under these two conditions, by the inverse function theorem (Lee [1], Theorem 4.14), $f_\theta|U$ is a $C^k$ immersion. Furthermore, if $U$ is compact and $f\theta|U$ is injective, then since a continuous injection from a compact space to a Hausdorff space is a homeomorphism (Lee [1], Theorem 4.25), $f\theta|U$ is a $C^k$ embedding, and $f\theta(U)$ is a $d$-dimensional $C^k$ embedded submanifold of $\mathbb{R}^n$.

Let us explain in detail what "full rank" means. For the Jacobian matrix $J \in \mathbb{R}^{n \times d}$ ($n > d$) to have full rank $d$ means that its $d$ column vectors are linearly independent, i.e., $\ker J = {0}$. An equivalent condition is that in the singular value decomposition $J = U\Sigma V^\top$, all $d$ singular values $\sigma_1 \geq \sigma_2 \geq \cdots \geq \sigma_d$ are positive ($\sigma_d > 0$).

The geometric intuition behind the full-rank condition is that "the decoder maps the $d$ independent directions in the latent space to $d$ independent directions in data space." If $\sigma_d = 0$ (rank less than $d$), then some direction in the latent space is "collapsed" by the decoder—moving in that direction produces no change in data space. In this case, the pullback metric $g_{ij} = (J^\top J)_{ij}$ degenerates (ceases to be positive definite), its inverse $g^{ij}$ does not exist, and Christoffel symbol computation becomes impossible.

(2) Vector Diffusion Maps (Singer & Wu, 2012).

Constructs a graph connection Laplacian from discrete data points ${z_i}{i=1}^N$. Finds $k$ neighbors for each data point, weights them by Gaussian kernel $w{ij} = \exp(-|z_i - z_j|^2/\varepsilon)$, and assigns an optimal rotation matrix $O_{ij} \in O(d)$ between local tangent spaces for each edge $(i,j)$. The graph connection Laplacian $S_\varepsilon$ is defined by

$$
(S_\varepsilon v)i = \frac{1}{d_i}\sum_j w{ij}(v_i - O_{ij}v_j), \quad d_i = \sum_j w_{ij}
$$

Continuous limit theorem (Singer & Wu, 2012). Under Assumption A1, as $N \to \infty$, $\varepsilon \to 0$, satisfying the rate condition $N\varepsilon^{d/2+1} \to \infty$, $S_\varepsilon$ converges in the spectral sense to the connection Laplacian on the tangent bundle of the continuous manifold $M_0$.

Note on assumptions. Assumption A1 as stated in this book is a simplified version. The full conditions required by Singer & Wu (2012) include additional technical requirements (reachability conditions on the manifold, relationships between kernel bandwidth and the manifold's radius of curvature, etc.). Readers intending rigorous application should consult the original paper for the complete set of conditions.

This theorem guarantees the mathematical justification for extracting continuous geometric structure from discrete data. However, this is an asymptotic result, and quantitative evaluation of approximation error for finite $N$ is generally difficult.

(3) Neural ODE (Chen et al., 2018).

Parameterizes the right-hand side of a differential equation by a neural network $f_\theta$:

$$
\frac{dz}{dt} = f_\theta(z, t), \quad z(0) = z_0
$$

Computes $z(0) \to z(T)$ via an ODE solver (e.g., Runge-Kutta), and learns the parameters $\theta$ through a loss function. Backpropagation uses the adjoint method for memory-efficient gradient computation.

The mathematical content of the adjoint method: the adjoint state $a(t) = \partial L / \partial z(t)$ ($L$ is the loss function) is obtained by solving the adjoint equation

$$
\frac{da}{dt} = -a(t)^\top \frac{\partial f_\theta}{\partial z}(z(t), t)
$$

backwards in time. This is a discrete analogue of the adjoint equation $\dot{p}_k = -\partial H / \partial z^k$ in Pontryagin's maximum principle (used in Chapter 13).

Extending Neural ODE to manifolds amounts to rewriting the ODE from Euclidean space $\mathbb{R}^d$ onto a Riemannian manifold. Specifically, when the VAE's latent space $U \subset \mathbb{R}^d$ is covered by a single chart (as is typically the case for VAE latent spaces), the coordinate expression $\dot{z}^k = f_\theta^k(z, t)$ remains valid, though attention must be paid to preserving metric structure (use of geometric integrators, etc.).

(4) Automatic differentiation for tensor computation via PyTorch/JAX.

PyTorch's autograd module and JAX's grad function dynamically build computation graphs and automatically apply the chain rule, accurately computing gradients of any computable function. This has made the full-chain computation from Riemannian metric $g_{ij}(z) = (J_{f_\theta}^\top J_{f_\theta}){ij}$ to Christoffel symbols $\Gamma^k{ij}$ to Riemann curvature tensor $R^l{}_{ijk}$ programmatically executable without hand calculation.

In particular, the create_graph=True option allows further differentiation of first-order derivative results (computing higher-order derivatives), which is indispensable for curvature tensor computation.

▶ For readers who want to understand the AI model foundations

Over the past few years, four technological breakthroughs have come together. Let us explain each so they can be understood with AI knowledge alone.

(1) VAE (Variational Autoencoder)—compressing data into a "smooth, low-dimensional space."

The VAE is a neural network model proposed by Kingma and Welling in 2014. The key difference from ordinary autoencoders (models that compress and reconstruct data) is that the encoder outputs a probability distribution.

In a standard autoencoder, the encoder outputs a single latent representation $z$ for an input $x$. In a VAE, the encoder outputs the mean $\mu$ and variance $\sigma^2$ of a normal distribution, and the latent representation is obtained as $z = \mu + \sigma \cdot \varepsilon$ ($\varepsilon$ is a random sample from a standard normal distribution).

The advantage of this "probabilistic" design is that the latent space is more likely to be a smooth space without "holes." In a standard autoencoder, meaningful latent representations exist only around points corresponding to training data, but in a VAE, the "spread" of the probability distribution smoothly fills the gaps between training data points.

The reason this book uses a VAE is that this "smoothness" is a prerequisite for manifold structure. Furthermore, by using smooth activation functions like $\tanh$ in the decoder and verifying that the Jacobian matrix is full rank, we guarantee that the VAE's output is a mathematically rigorous manifold.

(2) Vector Diffusion Maps (VDM)—extracting "geometric structure" from discrete data.

VDM is a method proposed by Singer and Wu in 2012. While the VAE "creates a smooth latent space," VDM "estimates the geometric structure of the manifold (tangent space orientations, parallel transport information) from a set of discrete data points."

The internal operation is as follows: (1) Compute distances between data points and weight them with a Gaussian kernel. (2) Estimate local tangent spaces (the "direction" spaces near each point) from neighboring data points. (3) Represent the "orientation mismatch" between adjacent tangent spaces as rotation matrices. (4) Assemble this information into a "connection Laplacian" matrix.

The eigenvectors of the connection Laplacian can be used as "coordinates" representing the global geometric structure of the manifold. Singer and Wu's theorem guarantees that as the number of data points grows to infinity, these discrete computational results converge to the exact geometric structure on the continuous manifold.

(3) Neural ODE—learning the "laws" of differential equations from data.

Neural ODE is a model proposed by Chen et al. in 2018. While conventional neural networks (such as ResNet) are "stacks of discrete layers," Neural ODE is formulated as a "continuous transformation."

Viewing a ResNet residual block $h_{t+1} = h_t + f_\theta(h_t, t)$ as the limit of making the step size approach zero yields the continuous differential equation $dz/dt = f_\theta(z, t)$. Neural ODE solves this differential equation with an ODE solver (numerical integration method) to compute "a continuous transformation from input $z(0)$ to output $z(T)$."

This book uses Neural ODE because: (1) it can learn the "laws of change" in business environments and international affairs from data, (2) adding control inputs (policies) to the differential equation enables control simulation, and (3) the adjoint method for efficient gradient computation corresponds mathematically to Pontryagin's maximum principle (optimal control theory).

(4) Automatic differentiation—executing tensor computations "without hand calculation."

PyTorch's autograd and JAX's grad are technologies where the computer automatically applies the chain rule to accurately compute derivatives of arbitrary functions.

Recall the chain rule for composite functions from high school: $\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx}$. Automatic differentiation applies this chain rule automatically over hundreds of stages. Since a neural network is essentially a chain of composite functions, automatic differentiation automatically computes the derivative of the output with respect to parameters.

In this book, automatic differentiation is used for the following "full-chain computation":

  1. Compute the Jacobian matrix $J_{f_\theta}$ of the VAE decoder $f_\theta$ via automatic differentiation → pullback metric $g_{ij} = (J^\top J)_{ij}$
  2. Compute partial derivatives $\partial g_{jl}/\partial z^i$ of the metric via automatic differentiation → Christoffel symbols $\Gamma^k_{ij}$
  3. Compute partial derivatives $\partial \Gamma^l_{ik}/\partial z^j$ of Christoffel symbols via automatic differentiation → Riemann curvature tensor $R^l{}_{ijk}$
  4. Compute partial derivatives of vector fields via automatic differentiation → Lie bracket $[V, W]$ and Lie derivative $\mathcal{L}_V g$

Step 3 requires further differentiating a quantity that is already a first derivative (the Christoffel symbols). In PyTorch, this requires the create_graph=True option. This is why ReLU activation functions are inadmissible and smooth activation functions like tanh are required—at the "bending points" of ReLU, second derivatives are undefined, causing step 3 to break down.

▶ For business leaders and policymakers

Over the past few years, three technological breakthroughs have come together, finally making it possible to "apply the same mathematics as physicists to social data."

Breakthrough 1: Technology to convert data into a "smooth topographic map" (VAE). Input your company's business data, and out comes a "topographic map of the business environment." This topographic map is fundamentally different from the usual UMAP or t-SNE scatter plots. While a scatter plot is merely "an arrangement of points," this topographic map is a computable map with rigorously defined distances, slopes, and curvatures. Imagine the difference between a car navigation 2D map and Google Earth's 3D terrain data.

Breakthrough 2: Technology to "navigate" on the topographic map (Neural ODE + control theory). "If we take this policy action, in which direction do we move on the topographic map?" "What is the shortest route to the target?" "Are there any danger zones (unstable regions) along that route?"—all computable.

Breakthrough 3: Technology to automatically execute all tensor computations (automatic differentiation). The metric, Christoffel symbols, curvature tensor, Lie derivative—30 years ago, a mathematician would have needed weeks of hand calculation for these. Now, PyTorch's automatic differentiation lets a computer compute them accurately in seconds. No mathematics specialist is needed for the computation itself (though domain expertise is essential for interpreting the results).

By combining these tools, it has become possible for the first time, in a mathematically justified form, to construct a differentiable manifold from social data and to perform differential, tensor, and control simulations at the same level as physicists.


5. What you will be able to do after reading this book

▶ For readers who want to understand the mathematical foundations

This book enables the following mathematical operations to be performed on social data.

(A) Quantification of nonlinear structure via scalar curvature $\mathrm{Scal}(z)$.

The scalar curvature $\mathrm{Scal}(z) = g^{ij}\mathrm{Ric}{ij}$ ($\mathrm{Ric}{ij} = R^l{}_{ilj}$) of a Riemannian manifold $(M, g)$ represents the "overall degree of curvature" at each point as a single scalar value. At points where $\mathrm{Scal}(z) > 0$, the volume of geodesic balls is smaller than in Euclidean space (geodesics converge); at points where $\mathrm{Scal}(z) < 0$, the volume is larger (geodesics diverge). In the data context, $\mathrm{Scal}(z) > 0$ corresponds to "a stable structure where neighboring states converge," and $\mathrm{Scal}(z) < 0$ corresponds to "an unstable structure where neighboring states disperse."

This is information that PCA cannot in principle provide. PCA is a projection onto a linear subspace, and curvature—an inherently nonlinear quantity—cannot even be defined within PCA's framework.

(B) Quantitative evaluation of structural deformation by policy via Lie derivative $\mathcal{L}_V g$.

Policy interventions and business strategies are modeled as vector fields $V \in \mathfrak{X}(M)$. The Lie derivative of the metric tensor $g$,

$$
(\mathcal{L}V g){ij} = \nabla_i V_j + \nabla_j V_i
$$

describes how the metric changes along the flow $\phi_t$ of the vector field $V$. Specifically, $\phi_t^* g = g + t(\mathcal{L}_V g) + O(t^2)$.

The Killing equation $\mathcal{L}_V g = 0$ means that $V$ generates an isometry—that is, $V$'s flow preserves the metric. When $\mathcal{L}_V g \neq 0$, each component of $(\mathcal{L}V g){ij}$ quantitatively shows "which direction of the metric is deformed by how much."

Traditional policy evaluation methods (DID, regression discontinuity design, etc.) measure "differences in outcomes" and do not measure "deformation of the geometric structure of data space." The Lie derivative provides qualitatively different information.

(C) Computation of optimal control paths via geodesic equations and Pontryagin's maximum principle.

The geodesic equation

$$
\ddot{\gamma}^k + \Gamma^k_{ij}\dot{\gamma}^i\dot{\gamma}^j = 0
$$

is the critical point (Euler-Lagrange equation from fixed-endpoint variation) of the energy functional $E[\gamma] = \frac{1}{2}\int g_{ij}\dot{\gamma}^i\dot{\gamma}^j,dt$, giving the "most natural" (energy-minimizing) path on the manifold.

When control inputs $u(t) \in U_{\mathrm{ad}} \subset \mathbb{R}^m$ exist, the affine control system

$$
\dot{z}^k = F^k(z) + \sum_{\alpha=1}^m u_\alpha(t) B^k_\alpha(z)
$$

yields, via Pontryagin's maximum principle, the Hamiltonian

$$
H(z, p, u) = p_k\left(F^k(z) + u_\alpha B^k_\alpha(z)\right) - \ell(z, u)
$$

and the canonical equations $\dot{z}^k = \partial H/\partial p_k$, $\dot{p}_k = -\partial H/\partial z^k$, from which the optimal control $u^(t)$ and optimal path $z^(t)$ are obtained.

(D) Interference detection for multiple policies via Lie bracket $[V_A, V_B]$.

The Lie bracket $[V_A, V_B]$ of two vector fields (two policies) $V_A, V_B$ measures the non-commutativity of the two flows. If $[V_A, V_B] = 0$, the flows commute (the result is independent of execution order); if $[V_A, V_B] \neq 0$, they do not commute (the result depends on order).

The magnitude $|[V_A, V_B]|$ quantitatively shows "the strength of interference between two policies," and the direction of $[V_A, V_B]$ shows "the direction of the displacement caused by interference."

▶ For readers who want to understand the AI model foundations

The four capabilities obtained from this book, explained from the AI implementation perspective:

(A) Curvature map generation. Compute scalar curvature at grid points on the VAE latent space and visualize as a heatmap. Positive curvature (red) = stable region, negative curvature (blue) = unstable region. This is information about "the quality of structure" absent from UMAP scatter plots. The full-chain computation from metric → Christoffel symbols → curvature tensor → scalar curvature is executed via PyTorch autograd.

(B) Policy impact simulation. Learn the policy vector field $V$ with a neural net (estimated from time-series data differences). Compute Lie derivative $\mathcal{L}_V g$ and evaluate how the policy changes the metric structure. $|\mathcal{L}_V g| \approx 0$ means structure-preserving; large means structure-deforming.

(C) Optimal path computation. Compute geodesics as numerical solutions via Neural ODE. Controlled optimal paths are computed via Pontryagin's maximum principle using the adjoint method (same mathematics as Neural ODE backpropagation).

(D) Policy interference detection. Compute Jacobian matrices of two vector fields with autograd and evaluate Lie bracket $[V_A, V_B] = JW \cdot V - JV \cdot W$. Close to zero means no interference; large means interference.

▶ For readers who want to apply this to business and policy decisions

Case Study A: Manufacturing Company X—Business Environment Manifold

Company X is a mid-sized manufacturer with approximately $1 billion in annual revenue. Its core domestic market has matured, and it is considering expansion into emerging markets. The strategy department must answer the following questions.

Question 1: How "different" is the emerging market's business environment from the current domestic market? This is a "distance problem." Using the Riemannian metric (Chapter 2), compute the Riemannian distance between the "domestic market business state" and the "emerging market business state." Euclidean distance (straight-line distance) and Riemannian distance (terrain-adjusted distance) can differ greatly, and the Riemannian distance more accurately reflects the "practical difficulty of transition."

Question 2: Does the transition path to the emerging market pass through a "stable route"? This is a "curvature problem." Compute the scalar curvature (Chapter 3) along the transition path and check whether it passes through "unstable regions" (regions of negative curvature). Routes passing through unstable regions carry a high risk of encountering sudden environmental changes.

Question 3: Which of several strategies provides the "lowest-risk path"? This is an "optimal path problem." Compute geodesics (Chapter 13) for the "most natural transition path" and use Pontryagin's maximum principle to compute the "optimal path under the constraints of available strategies."

Question 4: Do multiple strategies "interfere with each other" or "function independently"? This is a "Lie bracket problem." Compute the Lie bracket (Chapter 4) of the vector fields of two strategies to quantify the presence, direction, and magnitude of interference.

The dataset Company X should collect:

  • Company data: 50 locations × monthly (revenue, average transaction value, foot traffic, inventory turnover, headcount, customer satisfaction, etc.—15 indicators) × 10 years
  • Macroeconomic data: GDP growth rate, inflation rate, unemployment rate, consumer confidence index, exchange rate, interest rates, raw material price index
  • Industry data: market size, competitor shares, raw material prices, supply chain indicators, patent filings
  • Emerging market data: target country GDP, demographics, middle-class ratio, political stability index, regulatory environment index, infrastructure status

A total of approximately 30 indicators and approximately 6,000 data points (monthly × 10 years × 50 locations).

Case Study B: National Security Council—International Landscape Manifold

A national security council must determine the optimal combination of alliance policy and economic security policy in a rapidly changing international environment.

Question 1: Is the current international landscape "stable" or "unstable"? Answered by curvature analysis (Chapter 3).

Question 2: Does the alliance-strengthening policy "preserve the structure of the international order" or "change the structure"? Answered by Lie derivative (Chapter 4).

Question 3: If economic sanctions and alliance strengthening are executed simultaneously, do they "interfere"? Answered by Lie bracket (Chapter 4).

Question 4: What is the "minimum-risk path" to the target international order? Answered by optimal control (Chapter 13).

Intelligence dataset to collect:

  • Military indicators: each country's military spending/GDP ratio, nuclear warheads (count, delivery systems), conventional forces (active personnel, major equipment), cybersecurity capability index, space capability index
  • Diplomatic indicators: alliance network (bilateral alliance existence and strength), UN General Assembly voting patterns (voting concordance matrix), bilateral treaty count, diplomatic exchange count, summit frequency
  • Economic security indicators: trade dependency matrix (bilateral trade/GDP ratio), energy import dependency (fossil fuel import source concentration), rare earth/semiconductor strategic material procurement concentration, financial sanction vulnerability index (USD-denominated transaction dependency, SWIFT dependency)
  • Social/political indicators: democracy index (V-Dem, Freedom House), press freedom index (Reporters Without Borders), regime stability index, domestic public opinion on foreign policy support

193 countries × 40 indicators × 30 years = approximately 230,000 data points (annual data). Adding monthly trade/financial data brings this to over 700,000 points.

6. Conditions under which this method legitimately functions

▶ For readers who want to understand the mathematical foundations

The mathematical legitimacy of this procedure depends on four assumptions.

Assumption A1 (Manifold hypothesis). $X = {x_i} \subset \mathbb{R}^n$ is a noisy sample from a neighborhood of a $d$-dimensional compact connected $C^\infty$ submanifold $M_0$.

Assumption A2 (Decoder regularity). (a) $f_\theta \in C^k$ ($k \geq 3$). (b) $\mathrm{rank}, J_{f_\theta}(z) = d$ ($\forall z \in U$).

Assumption A3 (Sufficient data volume). $N\varepsilon^{d/2+1} \to \infty$.

Assumption A4 (Numerical stability). $\kappa(g(z))$ is bounded.

When assumptions break: A1 → manifold does not exist. A2(a) → curvature tensor undefined. A2(b) → metric degenerate. A3 → estimation inaccurate. A4 → numerical instability.

【Bridge from High School Mathematics】 Assumption A1 means "the data is concentrated near a low-dimensional surface." A2 means "the VAE decoder is a sufficiently good function." A3 means "there is enough data." A4 means "the computer's computations don't break down."

▶ For readers who want to understand the AI model foundations

The four conditions, explained in the language of AI models:

Condition 1 (Manifold hypothesis → data structure): The high-dimensional data actually lies on a low-dimensional "surface." The prerequisite for the VAE to learn successfully. Confirm with a PCA scree plot.

Condition 2 (Decoder regularity → model quality): The VAE decoder is a "good function." (a) Activation function choice: ReLU is kinked at $x=0$ and not differentiable ($C^0$). $\tanh$ is infinitely differentiable ($C^\infty$). ReLU cannot be used in this book. (b) Full-rank Jacobian: the decoder is "not collapsed." Different directions in the latent space map to different directions in data space. Verify with torch.linalg.svdvals.

Condition 3 (Data volume → learning quality): Sufficient data for VDM's continuous limit theorem to function. Rule of thumb: $N > 100 \times 2^d$.

Condition 4 (Numerical stability → computational reliability): The condition number of the metric tensor (max eigenvalue / min eigenvalue) is not too large. Regions with $\kappa > 10^4$ have unstable computations.

Critical warning: PyTorch will not produce errors even when conditions are not met. Curvature and Lie derivatives will be computed. But whether those numbers are meaningful depends on the conditions.

▶ For readers who want to apply this to business and policy decisions

This book's methods cannot build a manifold from just any data.

A computer executes computations as instructed. But making multi-billion-dollar investment decisions based on a "business environment manifold" constructed from inappropriate data is building on sand.

Data quality verification is just as important as the investment decision itself.


7. Five data quality checks

Check 1: Low-dimensional structure (Assumption A1). Does a PCA scree plot show 80%+ variance explained by a few principal components?

Case Study A: Company X, 30 indicators. 6 principal components explain 87%. → Pass.

Case Study B: 193 countries × 40 indicators. 8 principal components explain 82%. → Pass.

Check 2: Data density (Assumption A3). Guideline: $N > 100 \times 2^d$.

Case Study A: $d=6$, guideline 6,400 points. Monthly × 10 years = 6,000 points → borderline. Consider switching to weekly data.

Case Study B: $d=8$, guideline 25,600 points. Annual data = 5,790 points is insufficient → adding monthly data brings it to 70,000 points.

Check 3: Decoder fidelity (Assumption A2). Reconstruction error ≤ 5–10%. Jacobian minimum singular value $> 10^{-4}$.

Check 4: Stability of results (Assumption A4). Confirm stability by changing random seeds, varying $\varepsilon$, and bootstrapping.

Check 5: Consistency with other methods. Does the curvature pattern correspond to past known events?


8. Practical verification methods

▶ For readers who want to understand the mathematical foundations

A2(b) is SVD verification with $\sigma_{\min}(J_{f_\theta}(z)) > \delta$. A3 involves the $\varepsilon$-dependence of the spectral gap $\lambda_2 - \lambda_1$. A4 is monitoring $\kappa(g(z))$. A1 can only be verified indirectly via intrinsic dimension estimation (Facco et al. [6], Levina-Bickel [7]).

▶ For readers who want to understand the AI model foundations

Assumption Verification method Python tool
A1 Intrinsic dimension estimation + PCA scikit-dimension, sklearn
A2(a) Use $\tanh$/GELU/Softplus Activation function selection
A2(b) Jacobian SVD torch.linalg.svdvals
A3 Rule of thumb + bootstrap sklearn.utils.resample
A4 Condition number monitoring torch.linalg.cond

▶ For readers who want to apply this to business and policy decisions

Have your team confirm: "What is the reconstruction error?" "What is the minimum singular value of the Jacobian?" "Do results agree when the seed is changed?" If there are no clear answers, the analysis results cannot be trusted.


9. Practical recommendations

First, verify assumptions before and after analysis. Second, treat results as "exploratory insights." Third, use in conjunction with traditional methods. Fourth, always check robustness. Fifth, the geometric structures in this book are "nonlinear correlation structures"—not causal relationships.


Structure of this book

Part I (Chapters 1–4): Mathematical foundations. Topological spaces → differentiable manifolds → tangent spaces → Riemannian metrics → connections → curvature → Lie derivatives.

Part II (Chapters 5–8): Geometric AI tools. VAE → VDM → metric learning → Neural ODE → automatic differentiation tensor computation → the big picture.

Part III (Chapters 9–14): Practice. Data collection → manifold construction → geometric structure extraction → curvature analysis → Lie derivatives → control simulation → visualization and decision-making.

Each chapter features a three-layer structure (Mathematics / AI / Business), two case studies, quality checkpoints, and explicit cost-versus-return statements.

Now, let us begin the journey into the world of manifolds.

Chapter 1: From Topological Spaces to Differentiable Manifolds


What you gain from this chapter—return on investment

Chapter 1 is the "mathematical foundation work" for the entire book. Like a building's foundation, it is not the visible, glamorous part, but if this foundation crumbles, everything built upon it—curvature analysis, Lie derivatives, optimal control simulation—collapses. This chapter precisely defines "what a manifold is" and proves the conditions under which a VAE decoder's output becomes a manifold.

▶ For readers who want to understand the mathematical foundations: Why it's worth revisiting these basics

Proposition 1.1 (the manifold structure of the VAE decoder's image), established in this chapter, is the starting point for the mathematical legitimacy of the entire book. This proposition guarantees:

(i) The pullback metric $g_{ij}(z) = (J_{f_\theta}^\top J_{f_\theta}){ij}$ is a positive-definite Riemannian metric (foundation for Chapter 2).
(ii) The Christoffel symbols $\Gamma^k
{ij}$ are well-defined (foundation for Chapter 3).
(iii) The Riemann curvature tensor $R^l{}_{ijk}$ is well-defined (foundation for Chapter 3).
(iv) The Lie derivative $\mathcal{L}_V g$ computation is justified (foundation for Chapter 4).
(v) The geodesic equation and Pontryagin's maximum principle are applicable (foundation for Chapter 13).

All of (i)–(v) depend on the decoder $f_\theta$ being a $C^k$ immersion—specifically, $f_\theta \in C^k$ ($k \geq 3$) and the Jacobian matrix $J_{f_\theta}(z)$ being full rank. Furthermore, the upgrade from immersion to embedding requires compactness of the domain and the Hausdorff property of the codomain. Precisely understanding these topological conditions is essential for judging "why these conditions are needed" and "what breaks when they fail."

Cost: Somewhat abstract preparation, starting from the definition of topological spaces. Concepts like open sets, compactness, and the Hausdorff property may rarely be consciously considered by differential geometers who routinely perform tensor calculations. However, understanding what happens when these premises fail (for example, when the region of interest in the VAE latent space is not compact) requires the content of this chapter.

Return: Mathematical legitimacy for the next 13 chapters of computation. All of (i)–(v) above depend on Proposition 1.1, and omitting its proof leaves all subsequent computations in the state of "results are obtained but not mathematically justified."

▶ For readers who want to understand the AI model foundations: Why start from the definition of topological spaces

"Doesn't compressing with a VAE automatically give you a manifold?"—This intuition is half right, but half wrong.

A VAE decoder is a function from the latent space (a low-dimensional coordinate space) to data space (a high-dimensional space). Whether the image (the totality of outputs) of this function constitutes a "manifold" in the mathematically rigorous sense depends on the decoder's properties.

For example, the following problems can occur:

Problem 1: When the decoder is "collapsed." If the Jacobian matrix's rank drops (singular values are close to 0), moving in some direction in the latent space produces no change in data space. This means "one dimension of the latent space is wasted," and the decoder's image is not a $d$-dimensional manifold but a lower-dimensional structure. In this case, the pullback metric $g = J^\top J$ fails to be positive definite (has a zero eigenvalue), the inverse does not exist, and Christoffel symbols cannot be computed.

Problem 2: When the decoder is "kinked." A decoder using ReLU activation functions divides the input space into polyhedral regions, performing an affine (linear) mapping within each region. At region boundaries (hyperplanes where ReLU switches to 0), derivatives are discontinuous, and the second derivatives needed for curvature tensor computation are undefined.

Problem 3: When the decoder "self-intersects." If two different points $z_1 \neq z_2$ in the latent space are mapped to the same data point $f_\theta(z_1) = f_\theta(z_2)$, the decoder's image "intersects itself"—like a figure-eight. In this case, the topological structure of the decoder's image breaks down at the intersection point, and it fails to be a manifold.

This chapter mathematically rigorously defines the conditions under which none of these problems occur and proves that, under those conditions, the decoder's image is a manifold. To do so, we need to precisely define "what a manifold is," for which the concept of topological spaces is indispensable.

Cost: Time to understand the abstract concepts of topological spaces. "Open sets," "continuous maps," "homeomorphisms," "compactness," and the "Hausdorff property" are concepts that do not usually appear in AI/machine learning contexts.

Return: You will know precisely "what goes wrong in VAE design that breaks all subsequent computations." Specifically: why switching from ReLU to $\tanh$ is "necessary" ($C^k$ condition), why SVD verification of the Jacobian is "important" (full-rank condition), and why taking the "region of interest" in latent space as compact is "important" (immersion → embedding upgrade condition). All of these are essential for troubleshooting.

▶ For readers who want to apply this to business and policy decisions: Why foundation work is needed

This chapter corresponds to a building's "foundation work." Business leaders and policymakers need not directly understand the mathematical content of this chapter. However, whether the data science team understands this foundation decisively determines the reliability of analysis results.

An analogy: When you construct a new office building, you don't need to understand the details of the foundation work (pile depth, rebar arrangement, concrete mix). However, when the contractor says "we don't need piles," you should know this is a problem. Similarly, when the data science team says "ReLU is fine" or "we skipped Jacobian verification," you should know this is a serious problem in this book's approach.

What business leaders should verify: Ask your team, "Can you explain the four conditions for the VAE decoder to generate a manifold (compactness of the region of interest, smoothness, full rank, injectivity)?" A team that cannot clearly answer needs to study these foundations. A team that can clearly answer and show verification results provides grounds for trusting the analysis.


1.1 Topological Spaces and Continuous Maps

▶ For readers who want to understand the mathematical foundations

Definition of a Topological Space

Definition 1.1 (Topological space). A pair $(X, \mathcal{O})$ of a set $X$ and a family $\mathcal{O} \subset 2^X$ of its subsets is a topological space if $\mathcal{O}$ satisfies the following three conditions:

(T1) $\emptyset \in \mathcal{O}$ and $X \in \mathcal{O}$.

(T2) If ${U_\lambda}{\lambda \in \Lambda} \subset \mathcal{O}$ ($\Lambda$ is an arbitrary index set), then $\bigcup{\lambda \in \Lambda} U_\lambda \in \mathcal{O}$.

(T3) If $U_1, U_2, \ldots, U_n \in \mathcal{O}$ (finitely many), then $\bigcap_{k=1}^n U_k \in \mathcal{O}$.

$\mathcal{O}$ is called a topology on $X$, and elements of $\mathcal{O}$ are called open sets.

Remark 1.1. In (T2), the index set $\Lambda$ for the union is arbitrary (and may be uncountable), but in (T3), intersections are restricted to finitely many sets. This asymmetry is essential—the intersection of infinitely many open sets is not generally open. Example: On $\mathbb{R}$, $U_n = (-1/n, 1/n)$ is open for each $n$, but $\bigcap_{n=1}^\infty U_n = {0}$ is not open (in the standard topology on $\mathbb{R}$).

Example 1.1 (Standard topology on $\mathbb{R}^n$). Let $\mathcal{O}$ be the family of all sets expressible as arbitrary unions of open balls $B(x, r) = {y \in \mathbb{R}^n \mid |x - y| < r}$ in $\mathbb{R}^n$. Then $(\mathbb{R}^n, \mathcal{O})$ is a topological space. This is the standard (Euclidean) topology on $\mathbb{R}^n$, and throughout this book, when we write $\mathbb{R}^n$, we always assume this topology.

Example 1.2 (Discrete and indiscrete topologies). For any set $X$, $\mathcal{O} = 2^X$ (every subset is open) is the discrete topology, and $\mathcal{O} = {\emptyset, X}$ is the indiscrete topology. The discrete topology has "every subset open," while the indiscrete topology has "almost nothing open." Both are mathematically valid topologies, but in practice the standard topology on $\mathbb{R}^n$ is overwhelmingly important.

【Bridge from High School Mathematics】 The definition of a topological space is abstract, but essentially it establishes "the rules for which subsets we call 'open.'"

Consider the real line $\mathbb{R}$ from high school. The open interval $(0, 1) = {x \in \mathbb{R} \mid 0 < x < 1}$ is an "open set." Why? Because for any point in this interval, you can take a sufficiently small open interval around it (say $(0.4, 0.6)$) that still fits inside $(0, 1)$. In other words, every point in an open set has "room to move a little and still stay inside the set."

By contrast, the closed interval $[0, 1] = {x \in \mathbb{R} \mid 0 \leq x \leq 1}$ is not an open set. No matter how small an open interval you take around the endpoint $0$, it will contain negative numbers and extend beyond $[0, 1]$.

The three axioms (T1)–(T3) abstractly formalize this intuition of "having room." This allows us to define "nearness" and "continuity" not only in concrete spaces like $\mathbb{R}^n$ but also in general spaces like VAE latent spaces.

Continuous Maps and Homeomorphisms

Definition 1.2 (Continuous map). A map $f \colon (X, \mathcal{O}_X) \to (Y, \mathcal{O}_Y)$ between topological spaces is continuous if for every open set $V \in \mathcal{O}_Y$ in $Y$, its preimage $f^{-1}(V) = {x \in X \mid f(x) \in V}$ is an open set in $X$, i.e., $f^{-1}(V) \in \mathcal{O}_X$.

Remark 1.2. This definition generalizes the $\varepsilon$-$\delta$ definition of continuity. For the standard topology on $\mathbb{R}^n$, Definition 1.2 is equivalent to ordinary $\varepsilon$-$\delta$ continuity. The advantage of the topological definition is that continuity can be defined solely in terms of open set structure, without requiring a distance function (norm). This allows continuity to be handled even in spaces where distance is not naturally defined (such as certain function spaces or quotient spaces).

【Bridge from High School Mathematics】 This generalizes the high school definition of a continuous function—"as $x$ approaches $a$, $f(x)$ approaches $f(a)$." In the language of topological spaces, it says "pulling back a set of points considered 'near' in $Y$ (open set $V$) through $f$ yields a set of points considered 'near' in $X$ (open set $f^{-1}(V)$)."

Intuitively, it means "mapping nearby points to nearby points." If $f$ were not continuous, preimages of two nearby points in $Y$ might be scattered in $X$. Continuity guarantees this "preservation of nearness."

Definition 1.3 (Homeomorphism). A map $f \colon X \to Y$ between topological spaces is a homeomorphism if it satisfies all three conditions:

(H1) $f$ is surjective: for every $y \in Y$ there exists $x \in X$ with $f(x) = y$.
(H2) $f$ is injective: $f(x_1) = f(x_2)$ implies $x_1 = x_2$.
(H3) Both $f$ and $f^{-1}$ are continuous.

When a homeomorphism exists, $X$ and $Y$ are said to be homeomorphic, written $X \cong Y$.

Remark 1.3. Note the condition in (H3) that "$f^{-1}$ is also continuous"—a continuous bijection alone does not suffice. There exist continuous bijections whose inverses are not continuous (however, when the domain is compact and the codomain is Hausdorff, the inverse of a continuous bijection is automatically continuous—this is the essence of Theorem 1.1).

【Bridge from High School Mathematics】 A homeomorphism corresponds to the relation of "being deformable into one another like rubber."

For example, a square and a circle are homeomorphic. You can stretch a rubber square sheet into a circular shape. This deformation is continuous (no cutting or gluing), and it can be continuously reversed.

On the other hand, a sphere (the surface of a ball) and a torus (the surface of a doughnut) are not homeomorphic. A doughnut has a "hole," but a sphere does not. No continuous deformation of a sphere can create a hole. Creating a hole requires a "cut," and cutting is not continuous.

The study of such properties—invariant under continuous deformation, like "the number of holes"—is topology. This book does not delve into deep topological theory, but introduces the minimal topological concepts needed to define manifolds in this section.

Compactness

Definition 1.4 (Open cover and finite subcover). For a subset $A$ of a topological space $X$, a family ${U_\lambda}{\lambda \in \Lambda}$ of open sets is an open cover of $A$ if $A \subset \bigcup{\lambda \in \Lambda} U_\lambda$. A finite subfamily ${U_{\lambda_1}, \ldots, U_{\lambda_n}}$ satisfying $A \subset U_{\lambda_1} \cup \cdots \cup U_{\lambda_n}$ is called a finite subcover.

Definition 1.5 (Compactness). A topological space $X$ is compact if every open cover of $X$ has a finite subcover.

Theorem 1.0 (Heine-Borel theorem). A subset $K$ of $\mathbb{R}^n$ is compact if and only if $K$ is bounded and closed.

【Bridge from High School Mathematics】 The intuition for compactness is "fitting within a finite size and including the boundary."

For $\mathbb{R}^n$, by the Heine-Borel theorem, compact ⇔ bounded and closed.

  • Closed interval $[0, 1]$: bounded (between 0 and 1) and closed (includes endpoints 0 and 1) → compact.
  • Open interval $(0, 1)$: bounded but not closed (excludes endpoints) → not compact.
  • All of $\mathbb{R}$: closed but not bounded → not compact.

Compactness matters because compact sets often enjoy nice properties—e.g., "a continuous function on a compact set attains its maximum and minimum" (Weierstrass extreme value theorem). In this book, compactness is essential for the immersion → embedding upgrade theorem (Theorem 1.1).

Why compactness matters in the VAE context. The VAE latent space is all of $\mathbb{R}^d$, but the region where data is actually distributed is finite in extent (because KL regularization pushes the latent distribution close to a standard normal, so very few data points have $|z|$ large). If the "region of interest" $U$ is restricted to this data distribution range, $U$ can be taken as compact (bounded and closed). This satisfies the conditions for applying Theorem 1.1.

The Hausdorff Property

Definition 1.6 (Hausdorff property). A topological space $X$ is Hausdorff (also called a $T_2$ space) if for any two distinct points $p, q \in X$, there exist open sets $U, V \in \mathcal{O}$ with $p \in U$, $q \in V$, and $U \cap V = \emptyset$.

Example 1.3. $\mathbb{R}^n$ is Hausdorff. For $p \neq q$, take $r = |p - q|/2 > 0$, $U = B(p, r)$, $V = B(q, r)$; then $U \cap V = \emptyset$.

Remark 1.4. All spaces we commonly encounter ($\mathbb{R}^n$, sphere $S^n$, torus $T^n$, etc.) are Hausdorff. Non-Hausdorff spaces exist mathematically but are not encountered in this book. The Hausdorff property is important because it guarantees that "two distinct points can be topologically separated." Without it, open-set separation of two "distinct" points is impossible, and manifold structure breaks down.

【Bridge from High School Mathematics】 The Hausdorff property is a mathematically rigorous statement of the common-sense property that "two distinct points can be distinguished."

In the spaces we normally consider, it is obviously possible to take "non-overlapping" neighborhoods around two distinct points $p$ and $q$. For instance, around $p = 0$ and $q = 1$, we can take non-overlapping open intervals $(-0.3, 0.3)$ and $(0.7, 1.3)$.

Spaces where this "obvious property" fails are extremely "pathological" and do not arise in physics or data science. Nevertheless, for mathematical rigor, this condition is included in the definition of a manifold.

▶ For readers who want to understand the AI model foundations

The relationship between topological concepts and AI models. The mathematics of this section may seem unrelated to AI models, but it is indirectly important.

Continuous maps and neural networks. The VAE decoder $f_\theta$ is a continuous map. This is because the decoder is a composition of linear transformations (matrix multiplication + bias addition) and activation functions ($\tanh$, ReLU, etc.), and since linear transformations are continuous and $\tanh$ and ReLU are continuous, their composition is also continuous.

"The decoder is continuous" means "two nearby points $z_1$ and $z_2$ in latent space are mapped to two nearby points $f_\theta(z_1)$ and $f_\theta(z_2)$ in data space." This is the basis for "smoothness of the latent space." If the decoder were discontinuous, adjacent points in the latent space could map to completely different points in data space, and the "structure of the latent space" would cease to reflect the structure of data space.

Homeomorphisms and VAE quality. In an ideal VAE, the decoder learns a mapping "close to a homeomorphism." That is, it is desirable for the structures of the latent space and data space to be "topologically the same." However, if VAE training is insufficient or the latent dimension $d$ is chosen inappropriately, the decoder may exhibit "collapse" (rank drop) or "self-intersection," failing to be a homeomorphism.

Compactness and the VAE latent space. The VAE latent space is all of $\mathbb{R}^d$, but KL regularization causes the latent representations of training data to concentrate around the standard normal distribution $\mathcal{N}(0, I_d)$. In practice, 99.7% of the data lies within $|z| < 3\sigma$ ($\sigma = 1$), i.e., $|z| < 3$. Setting this region as the "region of interest" $U = {z \in \mathbb{R}^d \mid |z| \leq R}$ ($R$ an appropriate constant), $U$ is a bounded closed set and therefore compact by the Heine-Borel theorem.

▶ For readers who want to apply this to business and policy decisions

The mathematics in this section is preparation for precisely defining "manifold." Business leaders and policymakers do not need to directly understand it, but please remember one thing:

A "manifold" is "a smooth space that looks flat locally but is curved globally."

The Earth's surface is the most familiar example. Within a 100-meter radius of where you stand, the ground looks nearly flat. But the Earth as a whole is a sphere, not a plane. This property of "looking flat locally but being curved globally" is the essence of a manifold.

A company's business environment is similar. Within a specific market segment, the relationship between revenue and price might appear nearly linear. But viewed across multiple market segments, countries, and economic environments, the relationship curves nonlinearly. This book provides methods for mathematically constructing this "curved space" precisely and computing on it.


1.2 Differentiable Manifolds

▶ For readers who want to understand the mathematical foundations

Charts and Atlases

Definition 1.7 (Chart). A pair $(U, \varphi)$ of an open set $U$ of a topological space $M$ and a homeomorphism $\varphi \colon U \to \varphi(U) \subset \mathbb{R}^d$ is called a $d$-dimensional chart (coordinate neighborhood) on $M$. $\varphi$ is called the coordinate map, $\varphi(U)$ the image of the coordinate neighborhood, and $\varphi(p) = (z^1(p), \ldots, z^d(p))$ the local coordinates of the point $p$.

The essence of a chart is "mapping a part $U$ of the manifold $M$ to a part $\varphi(U)$ of $\mathbb{R}^d$ with 'the same shape' (homeomorphically)." This correspondence allows points on $M$ to be expressed in $\mathbb{R}^d$ coordinates $(z^1, \ldots, z^d)$, enabling ordinary calculus on $\mathbb{R}^d$.

【Bridge from High School Mathematics】 A chart is a "map." The Earth's surface (a sphere) is not flat, but a "map of the area around Japan" corresponds a part of the Earth to a plane (a sheet of paper). This correspondence is a "chart."

Maps always have "distortion," but for a sufficiently small area, the distortion is negligible. Mathematically, the chart $\varphi$ is a homeomorphism between $U$ and $\varphi(U)$, completely preserving the topological structure ("nearness" structure) of $U$. However, the metric structure (quantitative "distance" values) is not generally preserved. Discussion of metric structure takes place in Chapter 2.

Definition 1.8 ($C^k$ atlas). For an open cover ${U_\alpha}{\alpha \in A}$ of $M$ (i.e., $M = \bigcup\alpha U_\alpha$), a collection of charts

$$
\mathcal{A} = {(U_\alpha, \varphi_\alpha)}_{\alpha \in A}
$$

is a $C^k$ atlas if for all $\alpha, \beta \in A$ with $U_\alpha \cap U_\beta \neq \emptyset$, the transition function

$$
\varphi_\beta \circ \varphi_\alpha^{-1} \colon \varphi_\alpha(U_\alpha \cap U_\beta) \to \varphi_\beta(U_\alpha \cap U_\beta)
$$

is a $C^k$ map ($k$-times continuously differentiable).

Detailed explanation of transition functions. $U_\alpha \cap U_\beta$ is the "overlap" between charts $(U_\alpha, \varphi_\alpha)$ and $(U_\beta, \varphi_\beta)$. A point $p$ in this overlap can be expressed in both coordinate systems: $\varphi_\alpha(p) = (z^1_\alpha, \ldots, z^d_\alpha)$ and $\varphi_\beta(p) = (z^1_\beta, \ldots, z^d_\beta)$. The transition function $\varphi_\beta \circ \varphi_\alpha^{-1}$ provides the conversion rule from $\alpha$-coordinates to $\beta$-coordinates—a coordinate transformation.

"$C^k$" means this coordinate transformation is $k$-times differentiable with the $k$-th derivative also continuous. If coordinate transformations are smooth, calculus results in one coordinate system can be converted to another.

【Bridge from High School Mathematics】 An atlas is an "atlas of maps." A single map of the area around Japan cannot cover the entire Earth, so you collect maps of Europe, the Americas, Africa, etc. to cover the whole globe. This entire "book of maps" is the atlas.

Adjacent maps overlap. For example, a map of the Japan region and a map of the Russia region overlap in the Russian Far East. In this overlap, coordinate conversion is needed because the coordinates (how latitude and longitude are expressed) differ between the two maps. This coordinate conversion is the "transition function." If transition functions are smooth ($C^k$), calculus results join up consistently at map boundaries.

Definition 1.9 ($C^k$ differentiable manifold). A second-countable Hausdorff topological space $M$ is a $d$-dimensional $C^k$ differentiable manifold if $M$ has a maximal $C^k$ atlas.

Remark 1.5 (Second countability). Second-countable means the topology of $M$ is generated by countably many open sets. $\mathbb{R}^n$ is second-countable. This condition excludes pathological examples of manifolds (spaces with uncountably many connected components, etc.) and is always satisfied for the VAE latent spaces treated in this book.

Remark 1.6 (Maximal atlas). A $C^k$ atlas is not generally unique, but any $C^k$ atlas extends to a unique maximal $C^k$ atlas (the largest atlas containing all charts $C^k$-compatible with it). Therefore, specifying one $C^k$ atlas suffices to define a manifold, and the maximal atlas it generates uniquely determines the $C^k$ differentiable structure.

Remark 1.7 (Smoothness assumption in this book: $k \geq 3$). This book assumes $k \geq 3$. The reason is the following "chain of smoothness."

  • Step 6 (Chapter 3/Chapter 11): Computing the Riemann curvature tensor $R^l{}{ijk}$ requires first partial derivatives of the Christoffel symbols $\Gamma^k{ij}$.
  • Christoffel symbols are constructed from first partial derivatives of the Riemannian metric $g_{ij}$.
  • Therefore, computing $R^l{}{ijk}$ requires second partial derivatives of $g{ij}$, demanding $g_{ij} \in C^2$.
  • For the pullback metric (Chapter 2), $g_{ij}(z) = (J_{f_\theta}^\top J_{f_\theta}){ij}$, and when $f\theta \in C^k$, $J_{f_\theta} \in C^{k-1}$, hence $g_{ij} \in C^{k-1}$.
  • To obtain $g_{ij} \in C^2$, we need $k - 1 \geq 2$, i.e., $k \geq 3$.

Using $C^\infty$ activation functions ($\tanh$, GELU, Softplus) in all layers guarantees $f_\theta \in C^\infty$, amply satisfying the $k \geq 3$ condition. ReLU ($= \max(0, x)$) is $C^0 \setminus C^1$ (discontinuous derivative at $x = 0$), so not even $k \geq 1$ is guaranteed.


1.3 Immersions and Embeddings—Conditions for a VAE's Output to Be a Manifold

▶ For readers who want to understand the mathematical foundations

This section rigorously states and proves the conditions under which a VAE decoder's image becomes an embedded submanifold of $\mathbb{R}^n$. The key is the distinction between "immersion" and "embedding," and the theorem (Theorem 1.1) that a continuous injective immersion from a compact space is automatically an embedding.

Definition and Meaning of Immersion

Definition 1.10 (Immersion). A $C^k$ map $f \colon M \to N$ ($M$, $N$ are differentiable manifolds) is an immersion if at each point $p \in M$ the differential (tangent map)

$$
df_p \colon T_pM \to T_{f(p)}N
$$

is injective.

When $M$ is $d$-dimensional and $N$ is $n$-dimensional ($n \geq d$), using local coordinates, $df_p$ is represented by the Jacobian matrix $J_f(p) \in \mathbb{R}^{n \times d}$. The injectivity of $df_p$ is equivalent to $J_f(p)$ having rank $d$ (column full rank).

Intuitive meaning of immersion. An immersion is an operation of "mapping $M$ into $N$ without collapsing." "Without collapsing" means that all distinct directions on $M$ are mapped to distinct directions on $N$—the tangent map is injective.

Example 1.4 (Immersion but not embedding). Consider a map from $\mathbb{R}$ to $\mathbb{R}^2$ whose image forms a "figure eight." For instance, $\gamma(t) = (\sin 2t, \sin t)$ is regular (Jacobian full rank) on $(-\pi, \pi)$, but $\gamma(0) = \gamma(\pi) = (0, 0)$, creating a self-intersection. The tangent map is injective at each point (the immersion condition is satisfied), but $\gamma$ itself is not injective, so the image is not a manifold.

Definition and Meaning of Embedding

Definition 1.11 (Embedding). An immersion $f \colon M \to N$ is an embedding if $f$ is injective and $f \colon M \to f(M)$ (with $f(M)$ given the subspace topology from $N$) is a homeomorphism.

Meaning of embedding. An embedding is an operation of "mapping $M$ into $N$ without collapsing and without self-intersection." The image $f(M)$ of an embedding is a $d$-dimensional embedded submanifold of $N$.

Difference between immersion and embedding. An immersion requires "no collapse at each point's neighborhood" but does not require global properties (no self-intersection, the map to the image being a homeomorphism). An embedding additionally requires these global conditions.

Upgrade Theorem: From Immersion to Embedding

Theorem 1.1 (Lee [1], Theorem 4.25). If $M$ is a compact Hausdorff space and $N$ is a Hausdorff space, then a continuous injective immersion $f \colon M \to N$ is an embedding.

Proof sketch. $f$ is continuous and injective by assumption. We need to show $f \colon M \to f(M)$ is a homeomorphism, i.e., $f^{-1} \colon f(M) \to M$ is continuous. For a closed set $C \subset M$, since $M$ is compact and $C$ is closed, $C$ is compact. The image of a compact set under a continuous map is compact, so $f(C)$ is compact. Since $N$ is Hausdorff, compact sets are closed. Therefore $f(C)$ is closed in $N$, hence also closed in the subspace topology of $f(M)$. This means $f^{-1}$ is continuous (the preimage of a closed set being closed is equivalent to continuity). $\square$

Why this theorem is important. The immersion condition (Jacobian full rank) is a "local" condition, verifiable in a neighborhood of each point. The embedding condition is a "global" condition, requiring examination of the topological structure of the entire image. Theorem 1.1 shows that under the topological conditions of "compact Hausdorff domain," the local condition (immersion) + injectivity automatically implies the global condition (embedding). This reduces the verification that the VAE decoder's image is a manifold to the local condition (Jacobian full rank) + injectivity verification.

Manifold Structure of the VAE Decoder's Image

Proposition 1.1 (Manifold structure of the VAE decoder's image). Suppose the VAE decoder $f_\theta \colon U \to \mathbb{R}^n$ satisfies the following conditions:

(i) $U \subset \mathbb{R}^d$ is compact (a bounded closed set).
(ii) $f_\theta \in C^k$ ($k \geq 3$). Guaranteed by using $C^\infty$ activation functions ($\tanh$, GELU, Softplus).
(iii) $\mathrm{rank}, J_{f_\theta}(z) = d$ ($\forall z \in U$). That is, the Jacobian is full rank at every point in $U$.
(iv) $f_\theta|_U$ is injective.

Then $f_\theta(U)$ is a $d$-dimensional $C^k$ embedded submanifold of $\mathbb{R}^n$.

Proof.

Step 1 (Verifying immersion). Since $f_\theta \in C^k$ and $\mathrm{rank}, J_{f_\theta}(z) = d$ at each $z \in U$, by Definition 1.10, $f_\theta|_U$ is a $C^k$ immersion.

Here we explain why "Jacobian is full rank" is equivalent to "the differential (tangent map) is injective." The differential $df_{\theta,z} \colon T_z U \cong \mathbb{R}^d \to T_{f_\theta(z)}\mathbb{R}^n \cong \mathbb{R}^n$ is represented in local coordinates as the linear map $v \mapsto J_{f_\theta}(z) \cdot v$ by the Jacobian matrix $J_{f_\theta}(z) \in \mathbb{R}^{n \times d}$. This linear map being injective is equivalent to $\ker J_{f_\theta}(z) = {0}$, i.e., $J_{f_\theta}(z)$ having rank $d$ (the maximum column rank).

Step 2 (Upgrade to embedding). $U$ is a compact subset of $\mathbb{R}^d$ (assumption (i)). $\mathbb{R}^n$ is Hausdorff. $f_\theta|U$ is a continuous ($C^k \subset C^0$) injective (assumption (iv)) immersion (Step 1). By Theorem 1.1, $f\theta|_U$ is a $C^k$ embedding.

Step 3 (Embedded submanifold structure). Since $f_\theta|U$ is a $C^k$ embedding, its image $f\theta(U)$ is a $d$-dimensional $C^k$ embedded submanifold of $\mathbb{R}^n$ (Lee [1], Proposition 5.2). $\square$

Consequences when each condition of Proposition 1.1 fails:

(i) fails ($U$ non-compact): The upgrade from immersion to embedding is not guaranteed. Even as an immersion, the image may fail to be a manifold (self-intersection, topological breakdown of the image).

(ii) fails ($f_\theta$ not $C^k$, e.g., ReLU): While $C^1$ suffices for the immersion definition ("$C^1$ with injective tangent map"), the curvature tensor computation needed in this book requires $C^3$ or higher. ReLU does not even satisfy $C^1$, so the immersion condition is undefined at points on the ReLU kink surface.

(iii) fails (rank drops below $d$ at some point): At that point, the pullback metric $g_{ij} = (J^\top J)_{ij}$ has a zero eigenvalue and ceases to be positive definite. The inverse $g^{ij}$ does not exist, and Christoffel symbol computation is impossible.

(iv) fails ($f_\theta$ non-injective): The image self-intersects and fails to be a manifold. The meaning of metric and curvature at intersection points is lost.

▶ For readers who want to understand the AI model foundations

Understanding Proposition 1.1 in "AI model language."

Proposition 1.1 states "the four conditions for a VAE decoder's output to be a mathematically legitimate manifold."

Condition (i): The region of interest is compact (bounded and closed).

Implementation meaning: Instead of the entire VAE latent space $\mathbb{R}^d$, set a finite range where data is distributed (e.g., $|z| \leq 3$) as the "region of interest" $U$. In PyTorch, this is as simple as computing the norm of latent representations and excluding points outside the range.

Condition (ii): The decoder is $C^k$ ($k \geq 3$).

Implementation meaning: Use nn.Tanh(), nn.GELU(), or nn.Softplus() for all activation functions. Do not use nn.ReLU().

# ✗ Inadmissible (C^0, curvature computation impossible)
decoder = nn.Sequential(nn.Linear(d, 64), nn.ReLU(), nn.Linear(64, n))

# ✓ Recommended (C^∞, all higher-order derivatives computable)
decoder = nn.Sequential(nn.Linear(d, 64), nn.Tanh(), nn.Linear(64, n))

Condition (iii): Jacobian is full rank.

Implementation meaning: After training, perform SVD of the Jacobian at representative points (at least 50–100) in the latent space and confirm the minimum singular value is sufficiently far from 0.

import torch

def check_jacobian_rank(decoder, z_samples, latent_dim):
    """Verify the full-rank condition of the Jacobian matrix"""
    min_singular_values = []
    for z in z_samples:
        z = z.requires_grad_(True)
        f = decoder(z)
        n_out = f.shape[-1]

        # Compute the Jacobian matrix
        J = torch.zeros(n_out, latent_dim)
        for a in range(n_out):
            grad = torch.autograd.grad(f[a], z, retain_graph=True)[0]
            J[a] = grad

        # Singular value decomposition
        svs = torch.linalg.svdvals(J)
        min_sv = svs[-1].item()
        min_singular_values.append(min_sv)

    min_overall = min(min_singular_values)
    print(f"Minimum of minimum singular values: {min_overall:.6f}")
    if min_overall > 1e-4:
        print("✓ Full-rank condition satisfied (Assumption A2(b))")
    else:
        print("✗ WARNING: Rank deficiency detected at some points")
    return min_singular_values

Condition (iv): Decoder is injective.

Implementation meaning: For distinct latent points $z_1 \neq z_2$, $f_\theta(z_1) \neq f_\theta(z_2)$. Exact verification is difficult, but approximate verification is possible:

def check_injectivity(decoder, z_samples, threshold=1e-3):
    """Approximate verification of decoder injectivity"""
    outputs = [decoder(z).detach() for z in z_samples]
    n = len(outputs)
    collisions = 0
    for i in range(n):
        for j in range(i+1, n):
            if torch.norm(z_samples[i] - z_samples[j]) > threshold:
                if torch.norm(outputs[i] - outputs[j]) < threshold:
                    collisions += 1
                    print(f"  Latent points {i} and {j} map to the same data point")
    if collisions == 0:
        print("✓ Approximate injectivity verification passed")
    else:
        print(f"{collisions} collisions detected")

▶ For readers who want to apply this to business and policy decisions

Case Study A: Company X—Confirming the "topographic map is correct"

Company X's data science team verified the four conditions after VAE training.

Condition 1 (Compactness of region of interest): Set $|z| \leq 3$ as the region of interest, containing 99.7% of data in the latent space. → Automatically compact. ✓

Condition 2 (Smoothness): Used nn.Tanh() in the decoder. → $C^\infty$ guaranteed. ✓

Condition 3 (Full rank): Performed Jacobian SVD at 100 latent space points. Minimum of minimum singular values was $8.3 \times 10^{-3}$ ($> 10^{-4}$). → Full-rank condition met. ✓

Condition 4 (Injectivity): Approximate verification with 1,000 random pairs. No collisions. → Approximate injectivity passed. ✓

All four conditions passed. Company X's business environment manifold is a mathematically legitimate $C^\infty$ embedded submanifold. The metric, Christoffel symbols, curvature tensor, and Lie derivative computed on this manifold are all well-defined and can be expected to correctly reflect business environment structure.

Case Study B: International landscape—The problem of sparse data regions

The security analysis team's verification found that for some countries with limited data availability (particularly authoritarian regime countries), the Jacobian's minimum singular value tended to drop. Specifically, in the neighborhood of certain countries' data, the minimum singular value fell to $3.2 \times 10^{-5}$, below the full-rank threshold ($10^{-4}$).

The team made the following judgments:

(1) Major countries and regions (G20 nations + key allies): Data density is high and the full-rank condition is satisfied. Proceed with analysis.

(2) Smaller countries with sparse data: The full-rank condition is borderline. Present analysis results with a "confidence: medium" flag.

(3) Some countries with particularly sparse data: The full-rank condition is not met. Curvature and Lie derivative results in these regions are unreliable. Recommend additional intelligence data collection to the security council.

This is an important lesson. Manifold analysis is not "uniformly reliable across all regions." Reliability is high in regions with dense, high-quality data and decreases in regions with sparse data. Analysis results should always be accompanied by a "confidence map."


Quality Checkpoint for This Chapter—Should You Proceed or Go Back?

▶ For readers who want to understand the AI model foundations

□ Checklist (at the end of Chapter 1)

  1. Are activation functions $\tanh$/GELU/Softplus? → Yes: proceed / No: change and retrain
  2. Is the Jacobian minimum singular value $> 10^{-4}$? → Yes: proceed / Partially No: exclude that region or reduce $d$
  3. Is reconstruction error below 10% of data variance? → Yes: proceed / No: change architecture
  4. Has decoder injectivity been confirmed? → Yes: proceed / Collisions found: increase latent dimension $d$

If any fail: Proceeding to Chapter 2 and beyond will yield untrustworthy metric, curvature, and Lie derivative results.

▶ For readers who want to apply this to business and policy decisions

Ask your team: "Are all four conditions met?"

  • All four pass → proceed to next chapter
  • Any one fails → do not proceed. Direct your team to make design changes.

Next chapter preview: Chapter 2 defines "direction" (tangent vectors) and "distance" (Riemannian metric) on the manifold.

Case Study A preview: For the first time, "distance" is defined on Company X's business environment manifold, and "how far apart are the emerging market and the domestic market?" is quantified.

Case Study B preview: "Is Australia's security environment 'closer' to that of the US than Canada's?" can be answered quantitatively using Riemannian distance.

Chapter 2: Tangent Spaces and the Riemannian Metric


What you gain from this chapter—return on investment

Chapter 1 built the "manifold" stage. But a stage alone does nothing. This chapter defines two fundamental structures on the manifold. The first is the tangent space—defining "which directions you can move" at each point. The second is the Riemannian metric—enabling measurement of "distance" and "angle" on the manifold. With these two in place, the manifold transforms from a mere "shape" into a "topographic map" on which quantitative computation is possible.

▶ For readers who want to understand the mathematical foundations

The tangent space $T_pM$ is the space of "permitted velocity vectors" at each point $p$ of the manifold, forming a $d$-dimensional vector space. The tangent bundle $TM = \bigsqcup_{p \in M} T_pM$ is the domain for vector field definitions, indispensable for the Lie derivative in Chapter 4 and the control vector fields in Chapter 12.

The Riemannian metric $g$ is a family of smoothly varying positive-definite symmetric bilinear forms on each tangent space. Introducing the metric enables:

(i) Curve length $L[\gamma] = \int \sqrt{g_{ij}\dot{\gamma}^i\dot{\gamma}^j},dt$ and geodesic definition.
(ii) Quantification of VAE latent space geometric structure via the pullback metric $g_{ij}(z) = (J_{f_\theta}^\top J_{f_\theta}){ij}$.
(iii) Prerequisite for Christoffel symbol computation in Chapter 3 ($g
{ij} \in C^{k-1}$ required).
(iv) Scalar curvature computation $\mathrm{Scal}(z) = g^{ij}\mathrm{Ric}{ij}$ (Chapter 3).
(v) Lie derivative of the metric $(\mathcal{L}V g){ij} = \nabla_i V_j + \nabla_j V_i$ (Chapter 4).
(vi) Geodesic computation via energy functional minimization $E[\gamma] = \frac{1}{2}\int g
{ij}\dot{\gamma}^i\dot{\gamma}^j,dt$ (Chapter 13).
(vii) Hamiltonian formulation $H(z,p,u)$ in Pontryagin's maximum principle (Chapter 13).

In other words, the metric is the foundation of every computation from Chapter 2 onward. Without the metric, none of (i)–(vii) are defined.

2.1 Tangent Vectors and Tangent Spaces

▶ For readers who want to understand the mathematical foundations

Definition 2.1 (Tangent vector). For a smooth curve $\gamma \colon (-\varepsilon, \varepsilon) \to M$ ($\gamma(0) = p$, $\varepsilon > 0$) through a point $p$ of a $d$-dimensional differentiable manifold $M$, using a chart $(U, \varphi)$ ($p \in U$), the tangent vector of $\gamma$ at $p$ is

$$
v = \dot{\gamma}(0) = \left.\frac{d}{dt}\right|_{t=0} (\varphi \circ \gamma)(t) \in \mathbb{R}^d
$$

Definition 2.2 (Tangent space). The tangent space $T_pM$ at $p \in M$ is the set of all velocity vectors of smooth curves through $p$, forming a $d$-dimensional vector space. In local coordinates $(z^1, \ldots, z^d)$, the standard basis of $T_pM$ is ${\partial/\partial z^i|p}{i=1}^d$, and any tangent vector $v \in T_pM$ is uniquely expressed as $v = v^i \partial/\partial z^i|_p$ (Einstein summation convention).

Definition 2.3 (Tangent bundle). $TM = \bigsqcup_{p \in M} T_pM$ is the tangent bundle, a $2d$-dimensional differentiable manifold.

Definition 2.4 (Vector field). A vector field $V$ on $M$ is a smooth map $V \colon M \to TM$ (a smooth section of the tangent bundle, $V \in \Gamma(TM)$) assigning a tangent vector $V(p) \in T_pM$ smoothly to each point. In local coordinates: $V = V^i(z)\partial_i$.

A vector field is "an arrow placed at each point of the manifold." A weather map showing wind direction and speed at each location—arrows indicating the direction and strength of wind—is the most familiar example of a vector field on the Earth's surface ($S^2$).

2.2 Riemannian Metric—Defining "Distance" on the Manifold

▶ For readers who want to understand the mathematical foundations

Definition 2.5 (Riemannian metric). A Riemannian metric $g$ on a $d$-dimensional differentiable manifold $M$ is a family of positive-definite symmetric bilinear forms $g_p \colon T_pM \times T_pM \to \mathbb{R}$ on each tangent space, varying smoothly ($C^\infty$) with $p$. In local coordinates:

$$
g_p(u, v) = g_{ij}(z), u^i v^j
$$

where $g_{ij}(z)$ is a $d \times d$ positive-definite symmetric matrix-valued function.

【Bridge from High School Mathematics】 The simplest way to understand a Riemannian metric is as a "generalization of the Pythagorean theorem." The Pythagorean theorem gives $ds^2 = dx^2 + dy^2$ on the plane. A Riemannian metric generalizes this to $ds^2 = g_{ij},dz^i,dz^j$ where $g_{ij}$ varies from point to point—meaning "the way distance is measured differs at different locations."

Definition 2.6 (Curve length). The length of a smooth curve $\gamma \colon [a, b] \to M$ on $(M, g)$ is $L[\gamma] = \int_a^b \sqrt{g_{ij}(\gamma(t)), \dot{\gamma}^i(t), \dot{\gamma}^j(t)}, dt$.

Definition 2.7 (Energy functional). The energy of $\gamma$ is $E[\gamma] = \frac{1}{2}\int_a^b g_{ij}(\gamma(t)), \dot{\gamma}^i(t), \dot{\gamma}^j(t), dt$.

Theorem 2.1. Critical points of the energy functional (with fixed endpoints) are geodesics, satisfying $\ddot{\gamma}^k + \Gamma^k_{ij}\dot{\gamma}^i\dot{\gamma}^j = 0$.

Definition 2.8 (Pullback metric). For a VAE decoder $f_\theta \colon U \subset \mathbb{R}^d \to \mathbb{R}^n$ (satisfying Assumption A2), the pullback of the Euclidean metric on $\mathbb{R}^n$ is

$$
g_{ij}(z) = \sum_{a=1}^n \frac{\partial f_\theta^a}{\partial z^i}(z), \frac{\partial f_\theta^a}{\partial z^j}(z) = (J_{f_\theta}(z)^\top J_{f_\theta}(z))_{ij}
$$

Proposition 2.1. Under Assumption A2 ($\mathrm{rank}, J_{f_\theta}(z) = d$, $\forall z \in U$): (i) $g_{ij}(z)$ is positive-definite symmetric at each point. (ii) If $f_\theta \in C^k$, then $g_{ij} \in C^{k-1}$.

The pullback metric $g_{ij}(z)$ measures "how much distance change in data space is caused by an infinitesimal displacement in latent space." Directions where the decoder output is sensitive have "large Riemannian distance" (moving in that direction is "hard"); insensitive directions have "small distance" (moving is "easy").

▶ For readers who want to apply this to business and policy decisions

Case Study A: Company X—"Distance" between business states is defined for the first time

Business state pair Euclidean distance Riemannian distance Ratio
Domestic ↔ Southeast Asian market 2.3 5.8 2.5×
Domestic ↔ European market 3.1 2.1 0.7×
Domestic ↔ North American market 2.8 3.4 1.2×

The Southeast Asian market appears close by Euclidean distance but is far by Riemannian distance—a region of large data-space change (e.g., currency risk, regulatory differences) lies between them. Conversely, the European market appears far by Euclidean distance but is close by Riemannian distance—similar business environment structure.

Insight for leaders: Judging market "proximity" by Euclidean distance alone can lead to serious errors. The Southeast Asian market is "visually close but structurally costly to transition to"; the European market is "visually far but structurally easy to transition to."


Quality Checkpoint for This Chapter

  1. Is the pullback metric computable? Compute $g_{ij}(z)$ at 50+ representative latent-space points. NaN/Inf → check activation functions and create_graph=True.
  2. Is $g$ positive-definite at all points? Check all eigenvalues positive via torch.linalg.eigvalsh(g).
  3. Is the condition number ≤ $10^4$? Via torch.linalg.cond(g). Regions exceeding $10^4$ are numerically unstable.
  4. Does the metric reflect "meaningful structure"? Do equal-distance contours and $\det(g)$ spatial distribution align with domain knowledge?

Chapter 3: Connections, Covariant Derivatives, and the Curvature Tensor


What you gain from this chapter—return on investment

Chapter 2 defined "distance" on the manifold. This chapter introduces two more important structures. First, the covariant derivative—the method for "differentiating vectors" in curved space. Second, the curvature tensor—the quantity that numerically measures "how much space is curved." The scalar curvature $\mathrm{Scal}(z)$ derived from the curvature tensor is the core tool for creating "stability/instability maps" of business environments and international landscapes.

3.1 Connections and Christoffel Symbols

▶ For readers who want to understand the mathematical foundations

Definition 3.1 (Affine connection). An affine connection $\nabla$ on a manifold $M$ maps pairs of vector fields $(V, W)$ to a vector field $\nabla_V W$ satisfying: (C1) $C^\infty(M)$-linearity in the first variable, (C2) additivity in the second, (C3) Leibniz rule in the second.

Definition 3.2 (Christoffel symbols). $\nabla_{\partial_i}\partial_j = \Gamma^k_{ij}\partial_k$ defines the $d^3$ Christoffel symbols.

Theorem 3.1 (Uniqueness of Levi-Civita connection). On a Riemannian manifold $(M, g)$, there exists a unique affine connection satisfying: (LC1) torsion-free ($\Gamma^k_{ij} = \Gamma^k_{ji}$) and (LC2) metric-compatible ($\nabla_k g_{ij} = 0$). Its Christoffel symbols are:

$$
\Gamma^k_{ij}(z) = \frac{1}{2}g^{kl}(z)\left(\frac{\partial g_{jl}}{\partial z^i} + \frac{\partial g_{il}}{\partial z^j} - \frac{\partial g_{ij}}{\partial z^l}\right)
$$

The covariant derivative of a vector field $W = W^k\partial_k$ is $\nabla_{\partial_i} W = (\partial W^k/\partial z^i + \Gamma^k_{ij}W^j)\partial_k$—the first term is the ordinary partial derivative (component change rate), and the second is the correction for "apparent change due to the coordinate system's curvature."

3.2 Riemann Curvature Tensor and Scalar Curvature

Definition 3.3 (Riemann curvature tensor). When $\Gamma^k_{ij} \in C^1$ ($g_{ij} \in C^2$, $f_\theta \in C^3$ required):

$$
R^l{}{ijk}(z) = \frac{\partial\Gamma^l{ik}}{\partial z^j} - \frac{\partial\Gamma^l_{jk}}{\partial z^i} + \Gamma^l_{jm}\Gamma^m_{ik} - \Gamma^l_{im}\Gamma^m_{jk}
$$

The geometric meaning of the curvature tensor is path-dependence of parallel transport. Transporting a vector $v$ first by $\varepsilon$ along $\partial_i$ then by $\varepsilon$ along $\partial_j$, versus the reverse order, yields a difference of $\varepsilon^2 R(\partial_i, \partial_j)v + O(\varepsilon^3)$.

Definition 3.4 (Ricci tensor and scalar curvature).

$$
\mathrm{Ric}{ij} = R^l{}{ilj}, \quad \mathrm{Scal}(z) = g^{ij}(z),\mathrm{Ric}_{ij}(z)
$$

Convention note. The Ricci tensor is defined here by contracting the first and third indices of the Riemann curvature tensor: $\mathrm{Ric}{ij} = R^l{}{ilj}$. Some references (e.g., do Carmo's Riemannian Geometry) instead contract the first and fourth indices, $\mathrm{Ric}{ij} = R^l{}{ijl}$. Because $R^l{}{ijk} = -R^l{}{ikj}$, these two conventions differ by a sign: $R^l{}{ijl} = -R^l{}{ilj}$. This book consistently uses $\mathrm{Ric}{ij} = R^l{}{ilj}$ (the convention of Lee [1] and Misner-Thorne-Wheeler), so that $\mathrm{Scal} > 0$ corresponds to positively curved spaces such as the sphere. When comparing with other sources, please verify which convention is in use.

$\mathrm{Scal}(z) > 0$: geodesic balls have smaller volume than Euclidean (geodesics converge, stable). $\mathrm{Scal}(z) < 0$: larger volume (geodesics diverge, unstable).

▶ For readers who want to apply this to business and policy decisions

Curvature Terrain analogy Business meaning Strategic implication
$\mathrm{Scal} > 0$ (positive) Rolling hills Stable market structure Normal operations suffice
$\mathrm{Scal} \approx 0$ Flat plains Linear approximation valid Traditional regression works
$\mathrm{Scal} < 0$ (negative) Cliffs, steep slopes Structurally unstable region Strengthen risk management. Staged investment, pre-set exit conditions required

Case Study A: Company X—Scalar Curvature Map

The team computed a scalar curvature map of the business environment manifold and presented it as a color-coded topographic map to the executive committee.

Finding: The core domestic market region has positive scalar curvature (red, stable). The boundary between domestic and emerging markets has strongly negative curvature ($-2.3$, blue, unstable).

Validation: Business states during the 2008 financial crisis and COVID-19 (2020) were confirmed to lie in this negative-curvature region. The curvature analysis is consistent with past crises, supporting analytical reliability.

Insight for leaders: If the emerging market entry path passes through this unstable region, risk buffers (reserve funds, staged investment, pre-set exit conditions) are essential.


Chapter 4: Lie Derivatives and Killing Fields


What you gain from this chapter—return on investment

This chapter introduces the most distinctive and powerful tool in this book—the Lie derivative. The Lie derivative numerically measures "how a policy or strategy changes the structure of the space itself." This is something fundamentally impossible with any traditional data analysis method.

4.1 Flows and Lie Brackets

Definition 4.1 (Flow). The flow $\phi_t \colon M \to M$ of a vector field $V \in \mathfrak{X}(M)$ satisfies $\frac{d}{dt}\phi_t(p) = V(\phi_t(p))$, $\phi_0(p) = p$.

Definition 4.2 (Lie bracket). For $V, W \in \mathfrak{X}(M)$:

$$
[V, W]^k = V^j\frac{\partial W^k}{\partial z^j} - W^j\frac{\partial V^k}{\partial z^j}
$$

Geometric meaning: Flow by $V$ for $\varepsilon$, then by $W$ for $\varepsilon$, then reverse $V$ for $\varepsilon$, then reverse $W$ for $\varepsilon$. The displacement from the starting point is $\varepsilon^2[V,W] + O(\varepsilon^3)$. $[V,W] = 0$ ⇔ the two flows commute.

4.2 Lie Derivative of the Metric and the Killing Equation

Definition 4.3 (Lie derivative of the metric).

$$
(\mathcal{L}V g){ij} = \nabla_i V_j + \nabla_j V_i
$$

$\mathcal{L}_V g = 0$ is the Killing equation; $V$ satisfying it is a Killing field, generating isometries (metric-preserving transformations).

Proposition 4.1. $\mathcal{L}_V g = 0$ ⇔ $V$'s flow $\phi_t$ generates the isometry group ($\phi_t^* g = g$, $\forall t$).

▶ For readers who want to apply this to business and policy decisions

Case Study A: Company X—Structural impact of three strategies

Strategy $|\mathcal{L}_V g|$ Classification Meaning
Existing customer service enhancement 0.12 Structure-preserving Low risk / low return. Appropriate in stable periods.
Staged emerging market entry 1.8 Structure-deforming (moderate) Moderate risk / return. Manageable staged approach.
Fundamental portfolio restructuring 4.5 Structure-deforming (strong) High risk / high return. Requires thorough preparation.

Lie bracket analysis: $[V_{\mathrm{cost}}, V_{\mathrm{market}}]$ is large in the transition zone → "Do not pursue cost reduction and market expansion simultaneously. Separate the execution order."

Case Study B: Security policy structural impact

Alliance-strengthening policy $V_{\mathrm{alliance}}$: Lie derivative norm 1.2 (moderate deformation). Economic sanctions $V_{\mathrm{sanction}}$: 2.8 (strong deformation).

$[V_{\mathrm{sanction}}, V_{\mathrm{alliance}}]$ is large in certain tension states → "Strengthen alliances first, then impose sanctions" is optimal.


Part I complete. The four foundational concepts—manifolds (Chapter 1), metric (Chapter 2), connections/curvature (Chapter 3), Lie derivatives (Chapter 4)—are now in place.

Part II preview: Chapter 5 covers the VAE (Variational Autoencoder) in detail. Especially important for mathematicians unfamiliar with AI.

Part II: Tools of Geometric AI

Part I covered the four foundational concepts of differential geometry—manifolds, metrics, connections/curvature, and Lie derivatives. Part II covers the AI models needed to apply these mathematical concepts to social data: the VAE (Chapter 5), Vector Diffusion Maps and metric learning (Chapter 6), Neural ODE and automatic differentiation tensor computation (Chapter 7), and the big picture of Geometric Data Science (Chapter 8).

Readers strong in mathematics but weak in AI (including mathematicians) will find Part II most important. Conversely, readers strong in AI but weak in mathematics should return to Part I for the mathematical foundations.


Chapter 5: VAE and the Construction of Smooth Latent Manifolds


What you gain from this chapter—return on investment

The VAE (Variational Autoencoder) is the most important AI model supporting this book's entire approach. It compresses high-dimensional business or policy data into a low-dimensional "smooth space" and guarantees that this space is a mathematically legitimate manifold. This chapter is essential reading, especially for "mathematicians who know differential geometry but not VAEs."

▶ For readers who want to understand the mathematical foundations

A VAE is a probabilistic generative model that models the joint distribution $p_\theta(x, z) = p_\theta(x|z)p(z)$ of observed data $x \in \mathbb{R}^n$ and latent variables $z \in \mathbb{R}^d$ ($d \ll n$). The encoder $q_\phi(z|x)$ is the variational posterior, the decoder $p_\theta(x|z)$ is the likelihood, and both $\theta$ and $\phi$ are learned simultaneously by minimizing the ELBO (Evidence Lower Bound):

$$
\mathcal{L}(\theta, \phi; x) = -\mathbb{E}{q\phi(z|x)}[\log p_\theta(x|z)] + \mathrm{KL}(q_\phi(z|x) | p(z))
$$

The reason VAEs are critically important in this book's context is that the decoder $f_\theta : \mathbb{R}^d \to \mathbb{R}^n$ provides the constructing map for a differentiable manifold. As shown in Proposition 1.1 of Chapter 1, if $f_\theta \in C^k$ ($k \geq 3$), $\mathrm{rank}, J_{f_\theta}(z) = d$, and $f_\theta|U$ is injective, then $f\theta(U)$ is a $C^k$ embedded submanifold.

The most critical conditions for this book: (A) $f_\theta \in C^\infty$ (use $C^\infty$ activations: $\tanh$, GELU, Softplus in all layers; ReLU is inadmissible). (B) $\mathrm{rank}, J_{f_\theta}(z) = d$ ($\forall z \in U$), verified numerically after training.

5.2 Activation Functions and Smoothness—Why ReLU Cannot Be Used

Activation Definition Smoothness Usable? Reason
ReLU $\max(0, x)$ $C^0$ (derivative discontinuous at $x=0$) ✗ No Curvature tensor undefined
Leaky ReLU $\max(\alpha x, x)$ $C^0$ ✗ No Same as above
$\tanh$ $(e^x - e^{-x})/(e^x + e^{-x})$ $C^\infty$ ✓ Recommended All conditions met
GELU $x\Phi(x)$ ($\Phi$: standard normal CDF) $C^\infty$ ✓ Recommended Widely used in Transformers
Softplus $\log(1+e^x)$ $C^\infty$ ✓ Recommended Smooth approximation of ReLU
SiLU/Swish $x\sigma(x)$ ($\sigma$: sigmoid) $C^\infty$ ✓ Usable Similar to GELU
ELU piecewise $C^1$ △ Conditional $C^1$ but not $C^2$; insufficient for curvature

▶ For readers who want to apply this to business and policy decisions

Case Study A: Company X—Constructing the business environment manifold with VAE

Input data: 50 locations × monthly × 10 years, 30 indicators standardized. 6,000 data points × 30 dimensions.

VAE design: Input dim $n = 30$, latent dim $d = 6$, hidden layers 128 and 64, activation: all $\tanh$.

Results: Reconstruction error: 4.2% of data variance (< 10% threshold → ✓). Jacobian minimum singular value: $8.3 \times 10^{-3}$ (> $10^{-4}$ → ✓).


Chapter 6: Vector Diffusion Maps and Metric Learning


What you gain from this chapter

While the VAE (Chapter 5) "constructs the manifold," VDM (Vector Diffusion Maps) "extracts geometric structure from discrete data points." Additionally, this chapter covers pullback metrics and Cholesky parameterization for metric learning.

VDM (Singer & Wu, 2012) constructs a graph connection Laplacian $S_\varepsilon$ from discrete data. Continuous limit theorem: under Assumption A1, $S_\varepsilon$ converges spectrally to the connection Laplacian on the tangent bundle of the continuous manifold $M_0$.

VDM is "data-driven" and the pullback metric is "model-driven" geometric structure extraction. Comparing both for consistency increases analysis reliability.


Chapter 7: Neural ODE and Automatic Differentiation Tensor Computation


What you gain from this chapter

Neural ODE learns the "laws of change" in data; automatic differentiation automates the full-chain computation from metric → curvature → Lie derivative. Together, they make "policy simulation" and "tensor computation" executable in Python.

Neural ODE. The ODE $dz/dt = f_\theta(z, t)$ with neural net right-hand side $f_\theta$. Solved via ODE solver (Runge-Kutta etc.). Gradients via adjoint method—the discrete analogue of Pontryagin's maximum principle.

Automatic differentiation tensor computation. PyTorch autograd with create_graph=True enables: metric $g_{ij}$ → Christoffel symbols $\Gamma^k_{ij}$ → Riemann curvature tensor $R^l{}_{ijk}$ → Lie derivative $\mathcal{L}_V g$.


Chapter 8: The Big Picture of Geometric Data Science


What you gain from this chapter

This book's approach (manifold differential analysis) is just one part of the vast field called Geometric Data Science. This chapter surveys the field's landscape and clarifies where this book's approach sits.

Eight major approaches in Geometric Data Science:

  1. Manifold / Riemannian geometry ★This book's main approach — Pullback metrics, geodesics, curvature, control simulation
  2. Graph / network models — GCN, GAT, GraphSAGE
  3. Topological Data Analysis (TDA) — Persistent homology, Mapper algorithm
  4. Group / symmetry models — Equivariant networks, SE(3)-equivariant networks
  5. Non-Euclidean embedding models — Poincaré embeddings, hyperbolic space
  6. Geometric Deep Learning (unified framework) — Bronstein et al. (2021)
  7. Information geometry — Fisher information matrix, natural gradient descent
  8. Other geometric models — Morse theory, Ollivier-Ricci curvature, Sheaf Neural Networks

Which method should you use?

Data characteristics Best method Relation to this book
Many continuous indicators (business/economic) This book's methods Exactly this book
Network structure is key (supply chains, alliances) GNN (GCN, GAT) Complementary to this book's VDM
Want to investigate "holes" or "loops" TDA (persistent homology) Complementary to this book
Hierarchical data (org charts, taxonomies) Poincaré Embeddings Extensible from this book
Physical symmetries matter (molecules, 3D shapes) Equivariant NN Same mathematical spirit as Killing fields

Part II complete. The four tools of geometric AI and the Geometric Data Science landscape are now in place.


Part III: Constructing and Analyzing Manifolds from Social Data

Part I covered the mathematical foundations, Part II the geometric AI tools. Part III puts it all into practice. Case Study A (Company X's business environment manifold) and Case Study B (international landscape manifold) are executed with concrete Python code.


Chapter 9: Data Collection and Manifold Construction via VAE (Steps 1–2)


What you gain from this chapter

This chapter is the starting point of practice. It executes the first steps—data collection/preprocessing (Step 1) and VAE-based latent manifold construction (Step 2). The quality of these two steps determines the reliability of all subsequent analysis.

Case Study A: Company X's dataset

Company data (50 locations × monthly × 10 years): Revenue, average transaction value, foot traffic, inventory turnover, headcount, customer satisfaction, repeat rate, customer demographics, product category mix, store area, operating hours, advertising spend (15 indicators).

Macroeconomic data (monthly × 10 years): GDP growth, inflation (CPI), unemployment, consumer confidence, exchange rate (EUR/USD), interest rate (10-year Treasury yield), oil price, industrial production index (8 indicators).

Industry data (monthly × 10 years): Market size estimate, top 3 competitors' estimated revenue, raw material price index, supply chain delay index, patent filings, industry average profit margin (7 indicators).

Total: 30 indicators × 6,000 data points.

Quality Checkpoint (most important in the entire book)

  1. PCA confirms intrinsic dimension?
  2. Data volume meets $N > 100 \times 2^d$?
  3. VAE reconstruction error < 10% of variance?
  4. Jacobian minimum singular value > $10^{-4}$?
  5. Activation functions are $\tanh$/GELU/Softplus?

If any fail, do not proceed.


Chapter 10: Geometric Structure Extraction and Riemannian Metric Learning (Steps 3–4)

Step 3 applies VDM to the VAE latent space; Step 4 computes the pullback metric $g_{ij}(z) = (J_{f_\theta}^\top J_{f_\theta})_{ij}$ at lattice points. "Distance" is now defined, enabling quantitative comparisons between business states.


Chapter 11: Curvature Tensor Computation and Structural Analysis (Steps 5–6)

Step 5 computes Christoffel symbols $\Gamma^k_{ij}$ via automatic differentiation. Step 6 computes the Riemann curvature tensor $R^l{}{ijk}$, Ricci tensor $\mathrm{Ric}{ij}$, and scalar curvature $\mathrm{Scal} = g^{ij}\mathrm{Ric}_{ij}$. The scalar curvature map visualizes "stable regions" and "unstable regions."

Case Study A finding: Core domestic region has positive curvature (stable). Boundary with emerging markets: scalar curvature $-2.3$ (strongly unstable). Past crisis states (2008, 2020) confirmed in this region.


Chapter 12: Lie Derivative Analysis of Policy and Strategy Impact (Step 7)

This chapter is the climax for business leaders and policymakers. The Lie derivative computation shows "whether a strategy preserves or deforms the environment's structure" and "whether two strategies interfere."

$(\mathcal{L}V g){ij} = \nabla_i V_j + \nabla_j V_i$. $|\mathcal{L}_V g| \approx 0$ → structure-preserving. Large → structure-deforming. Lie bracket $[V_A, V_B]$ for interference detection.


Chapter 13: Control Simulation and Optimal Path Computation (Steps 8–9)

Step 8 formulates the affine control system on the manifold:

$$
\dot{z}^k(t) = F^k(z(t)) + \sum_{\alpha=1}^m u_\alpha(t) B^k_\alpha(z(t))
$$

Step 9 computes geodesics and optimal control paths via Pontryagin's maximum principle.

Case Study A: Company X—Optimal transition path to emerging markets

Path 1 (Linear plan): Shortest by Euclidean distance. Passes through negative-curvature unstable region. High risk.

Path 2 (Geodesic): Shortest by Riemannian distance. Detours around unstable region. Lowest energy.

Path 3 (Controlled optimal path): Under constraints of 3 available strategies (service enhancement, staged entry, joint venture), minimizes transition cost via Pontryagin's maximum principle. Yields optimal schedule: "First 6 months: service enhancement at max intensity → next 12 months: staged entry at moderate intensity → final 6 months: joint venture at high intensity."

Recommendation: Path 3—lowest risk with a concrete strategy schedule.


Chapter 14: Visualization, Interpretation, and Decision-Making (Step 10)

Integrates, visualizes, and connects to decision-making all results from Chapters 9–13. This is the goal of the book.

Distinction from causal inference: The geometric structures in this procedure are nonlinear correlation structures, not causal relationships. Causal inference requires intervention data, structural causal models, or natural experiments separately.

Error accumulation: Step 2 decoder error → Step 3 discretization error → Step 4 metric error propagate downstream. Curvature, involving second derivatives, is especially sensitive to error. Robustness checks are indispensable.

Case Study A: Company X—Report to the Executive Committee

Document 1: Business environment topographic map (scalar curvature map). Stable (red) and unstable (blue) regions. Past crises plotted.

Document 2: Lie derivative analysis. Structural impact comparison of 3 strategies. Interference analysis (Lie bracket) of cost reduction and market expansion.

Document 3: Optimal transition path. Comparison of 3 paths (linear, geodesic, controlled optimal). Recommended path with strategy schedule.

Document 4: Confidence map. Dense-data regions (high confidence) and sparse-data regions (low confidence). Explicitly stating "this analysis is not uniformly reliable across all regions."

Executive decision: "Adopt Path 3 (controlled optimal). Secure $3M reserve for passage through unstable region. Based on Lie bracket analysis, execute cost reduction and market expansion sequentially, not in parallel."

Final Quality Checkpoint

  1. Does the curvature map align with past known events?
  2. Does the Lie derivative ordering align with domain knowledge?
  3. Are optimal paths physically reasonable (no divergence, sensible range)?
  4. Are results stable under random seed changes and 10% data removal?
  5. Are results presented as "exploratory insights" (not definitive conclusions)?
  6. Is the distinction between correlation and causation clearly stated?

All Yes → analysis results can be used as reference for decision-making. Any No → return to the relevant chapter and redo.


Appendix A: List of Assumptions and Mathematical Justification

  • A1. Manifold hypothesis: Data is sampled from a neighborhood of a low-dimensional $C^\infty$ compact connected submanifold.
  • A2. Decoder regularity: (a) $f_\theta \in C^k$ ($k \geq 3$). (b) $\mathrm{rank}, J_{f_\theta}(z) = d$ ($\forall z \in U$).
  • A3. Sufficient data volume: Rate condition $N\varepsilon^{d/2+1} \to \infty$.
  • A4. Numerical stability: $\kappa(g(z))$ bounded.
Assumption Impact when violated
A1 Manifold itself does not exist. Entire procedure meaningless.
A2(a) Curvature tensor undefined (when using ReLU).
A2(b) Metric degenerates; inverse does not exist.
A3 VDM geometric structure extraction inaccurate.
A4 Tensor computations numerically unstable.

Appendix B: Python Environment Setup

# Basic environment
pip install numpy pandas scikit-learn matplotlib plotly

# Deep learning
pip install torch torchvision torchdiffeq

# TDA (Topological Data Analysis, used in Chapter 8)
pip install ripser persim giotto-tda

# Intrinsic dimension estimation (used in Check 1)
pip install scikit-dimension

# Visualization
pip install plotly jupyter

Appendix C: References

  1. J. M. Lee, Introduction to Smooth Manifolds, 2nd ed., Springer, 2012.
  2. A. Singer and H.-T. Wu, "Vector Diffusion Maps and the Connection Laplacian," Comm. Pure Appl. Math., 65(8), 2012.
  3. D. P. Kingma and M. Welling, "Auto-Encoding Variational Bayes," Proc. ICLR, 2014.
  4. R. T. Q. Chen et al., "Neural Ordinary Differential Equations," Proc. NeurIPS, 2018.
  5. G. Arvanitidis et al., "Latent Space Oddity," Proc. ICLR, 2018.
  6. E. Facco et al., "Estimating the intrinsic dimension of datasets," Scientific Reports, 7, 2017.
  7. E. Levina and P. J. Bickel, "Maximum likelihood estimation of intrinsic dimension," Proc. NeurIPS, 2004.
  8. L. S. Pontryagin et al., The Mathematical Theory of Optimal Processes, 1962.
  9. M. M. Bronstein et al., "Geometric Deep Learning," arXiv:2104.13478, 2021.
  10. M. Gidea and Y. Katz, "Topological Data Analysis of Financial Time Series," Frontiers in Applied Mathematics and Statistics, 2018.
  11. J. Hansen and T. Gebhart, "Sheaf Neural Networks," NeurIPS Workshop, 2020.
  12. C. Bodnar et al., "Neural Sheaf Diffusion," Proc. ICLR, 2022.
  13. S. Amari, Information Geometry and Its Applications, Springer, 2016.
1
1
1

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
1
1

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?