Forward-deployed / Learning zone
Knowledge graphsa standalone module
Lesson 05

Reasoning & analytics

TL;DR

The payoff for all that construction: once knowledge is a graph, whole families of questions become computable that tables and documents can't express at any price. Five families cover nearly everything on a product roadmap. Traversal follows chains: exposure, lineage, "how are these two connected?" Centrality finds which nodes matter most — the PageRank family. Community detection finds which nodes cluster — fraud rings, customer segments, duplicate candidates. Link prediction finds which edges are missing or likely next — recommendations, next-best-action. Rule-based inference derives facts that follow logically from stated facts — "subsidiary of a sanctioned entity is sanctioned-adjacent." Each family maps to a product capability you can name in a roadmap review, and each has a different trust profile. Traversal and inference produce explainable answers you can show a regulator. Predictions produce probabilistic ones you should ship as suggestions, not facts. Knowing which family a proposed feature sits in tells you its cost, its latency, and how much explanation your UX owes the user.

🎯 For the product leader

Why it matters — "Graph analytics" sounds like a data-science luxury. It's actually the menu of what your graph investment can ship: the fraud feature, the recommendation engine, the risk-exposure dashboard are all entries on this menu. If you can't name the families, you can't spot which roadmap items the graph makes cheap. You also can't be appropriately suspicious when a pitch calls a prediction a fact.

What it changes in your decisions — Roadmap items get tagged with their family, which sets three things at once: compute pattern (real-time query vs. batch job), explainability (deterministic vs. probabilistic), and eval burden (correctness vs. precision/recall on a golden set).

Ask yourself — "For this graph-powered feature: is the answer traversed, inferred, or predicted — and does the UX honestly reflect which one it is?"

Risk if ignored — A link prediction ships rendered like a stored fact. A customer asks "why does your product say we work with X?" The honest answer — "an algorithm guessed" — lands in a complaint, or a headline.

The five families, as product capabilities

Five families, three trust levels

What becomes computable · each family maps to a roadmap capability

Traversal and inference are explainable. Predictions are probabilistic and should ship as suggestions.

The knowledge graph
Explainable — a path or a rule
Traversal & paths

"How is A connected to B?" → risk exposure, lineage, 360° views

Rule-based inference

"What follows from the facts?" → compliance screening, entitlements

Structural — a score over the whole graph
Centrality

"Which nodes matter most?" → key accounts, single points of failure

Community detection

"What clusters exist?" → fraud rings, segmentation, duplicates

Predictive — a probability
Link prediction & embeddings

"Which edge is missing / next?" → recommendations, next-best-action

Traversal is the workhorse — the multi-hop questions from lesson 1, plus the underrated pathfinding form: "show every chain connecting this customer to this sanctioned entity." The answer is the path itself — evidence you can put in front of an auditor. It runs live, in the request path, if depth-capped.

Centrality ranks importance by structure. PageRank and its relatives find the supplier whose failure cascades furthest, the person the org actually routes through (rarely the org chart's answer), and the account whose loss would unravel a network. This is batch computation, refreshed on a schedule, and consumed as a score.

Community detection finds dense clusters no one labeled: accounts that share devices, addresses, and payment methods (a fraud ring, structurally, before any individual account misbehaves — the reason graph analytics is standard kit in financial crime), customers that cluster by actual usage rather than firmographics, and — a nice dogfooding touch — probable duplicates the resolution pipeline missed.

Link prediction guesses edges that are missing or coming: customers who look like this one also need that (recommendations), this person likely knows that person, this record is probably the same entity as that one. The modern machinery is graph embeddings — compressing each node's neighborhood into a vector so "structurally similar" becomes computable — and increasingly graph neural networks. It is powerful, and probabilistic to the bone: these are suggestions wearing math.

Rule-based inference derives facts logically: part-of chains roll up, ownership percolates through corporate trees, "handles EU personal data" propagates to every system downstream of one that does. It is deterministic and auditable, and it is the reason formal ontologies exist. In regulated domains, this family is the product.

The trust gradient

The families differ most in what kind of answer they produce — and your UX, evals, and compliance posture must track it:

Family Answer type Show the user Eval discipline
Traversal Deterministic path The path itself — it is the explanation Correctness + freshness of underlying facts
Inference Deterministic derivation The rule and the facts it fired on Rule review + regression tests, like code
Centrality / community Structural score Rank or grouping, framed as analysis Stability across runs; sanity panels with domain experts
Link prediction Probability A suggestion, with a why-shown ("shares 3 suppliers with...") Precision/recall on golden sets, online acceptance rates

The gradient also sets where mistakes hurt. A wrong stored fact corrupts every family downstream — this is why construction quality outranks algorithm choice. A wrong prediction, honestly framed as a suggestion, costs a shrug. The catastrophic combination is a prediction stored back into the graph as a fact: the guess launders itself into evidence, and six months later nobody remembers it was a guess. If predictions must be persisted, they must carry their confidence and provenance forever.

Operational shape

Two compute patterns, two cost profiles. Query-time work (traversal, inference on demand) lives in the request path: milliseconds, depth caps, timeouts. Feature latency is graph latency. Batch work (centrality, communities, embeddings, prediction scoring) runs on schedules and ships scores. It's cheap to serve, but stale by design — the fraud ring detected nightly is invisible for up to 24 hours, and whether that's fine is a product decision, not an infrastructure one. Most platforms bundle these algorithms into libraries, so the engineering lift is usually integration, not invention. The real work is choosing thresholds and owning what the scores mean.

Failure modes

Practitioner checklist