Knowledge Graphs
Facts as a network instead of a table: nodes, edges, and labels, why Neo4j models data this way, and why the same idea already decides what Google, and YouTube, show you next.
What a knowledge graph is
A knowledge graph is a semantic network, a network of real-world entities: events, situations, concepts, people, places. What makes it a graph rather than just a list is that it illustrates the relationship between those entities, not just names them.
Take two plain sentences: "Rohit Sharma is the captain of the Indian cricket team" and "Rishabh Pant is the wicketkeeper of the Indian cricket team." Pulling entities out of raw text like this, at scale, is its own NLP task called named entity recognition (NER), and it is usually the first step in building a graph automatically instead of by hand.
Once the entities are extracted, each sentence becomes a triple: (Rohit Sharma) -[CAPTAIN_OF]-> (Indian Cricket Team), (Rishabh Pant) -[WICKETKEEPER_OF]-> (Indian Cricket Team). Both triples share the same team node, and that shared node is where the interesting part starts.
Nodes, edges, and labels
A knowledge graph breaks down into exactly three components. Nodes are any object, place, or person, drawn as a circle. Edges are the relationship between two nodes, drawn as a line or arrow. Labels name what a node actually is, Person, Team, and so on, sitting underneath the entity's name.
With both triples sharing the Indian Cricket Team node, a relationship neither sentence stated appears on its own: Rohit Sharma and Rishabh Pant are teammates. Nobody wrote that sentence, it falls out of the graph's structure once two people connect to the same team.
Now add more text. A different article calls Virat Kohli and Rohit Sharma "cricket brothers." A third calls Hardik Pandya and Rohit Sharma close friends from their Mumbai Indians days. Each new sentence adds a node and an edge, and the graph never really finishes, it keeps extending across every page on the internet that mentions any of these people.
You've already been using one
Albert Einstein
Born 14 March 1879 · Died 18 April 1955
This is already running under every major search engine. Search "Albert Einstein" on Google and the panel that appears, birth date, death date, field, occupation, is built from exactly this structure: nodes and properties pulled from a graph, not copied off a single webpage.
Scroll down and there is a "people also search for" row: Isaac Newton, Stephen Hawking, J. Robert Oppenheimer. Those names show up because they are graph neighbors of Einstein, physicists from overlapping eras whose pages keep getting mentioned near his across the web, not because anyone typed their names into a search box.
The same graph decides what a YouTube video gets suggested next to, and which ads follow a search around afterward. Search for one data science YouTuber and the same mechanism surfaces others in that same neighborhood of the graph.
Property graphs vs. relational tables
A relational database, MySQL, SQL Server, models the same information as tables: rows for records, columns for fields, and constraints, primary keys, foreign keys, candidate keys, to keep it all consistent. Finding out that Rohit Sharma and Rishabh Pant are teammates means joining a players table to a teams table and matching on a shared team ID.
A property graph skips the join. Nodes hold data directly, and every node carries a label (its type) plus properties as key-value pairs, name: "Rohit Sharma", born: 1987. A relationship connects two nodes directly, so "teammates" is one edge walked, not a query planned around a join.
This is why graph databases fit entity-heavy, deeply interconnected data especially well: social graphs, recommendation systems, fraud rings, retrieval-augmented generation. Anywhere the real question is what's connected to this, a graph answers it by walking an edge instead of scanning and matching rows.
What a graph can compute that a table can't
Storing data as a graph isn't only about avoiding joins, it also unlocks algorithms that have no clean equivalent over rows and columns. Shortest path finds the fewest hops connecting two nodes, the same idea behind route planning or figuring out how two people are connected through mutual contacts.
Centrality measures rank nodes by influence rather than alphabetically or by timestamp: the node sitting on the most paths between everyone else floats to the top. This is how a fraud investigation finds the one account at the center of a web of suspicious transactions instead of checking every account by hand.
Community detection goes a step further and finds clusters of densely connected nodes automatically, groups that talk to each other far more than they talk to the rest of the graph. That's the same idea GraphRAG leans on to organize a document set into themes, and it shows up in drug discovery too, clustering diseases, symptoms, genes, and proteins by how tightly they interact, and even in personal study: a mind map of a subject's concepts is a knowledge graph, read by hand.
Neo4j and the property graph model
Neo4j is the graph database most of this ecosystem is built around, chosen for real-time insight into how entities connect and a query language built specifically to make that retrieval easy. Its underlying data model is exactly the property graph just described: nodes, relationships, and properties.
Every node gets a label, its type, and any number of properties as key-value pairs. A Person node might carry name and born; nothing forces every node of the same label to carry identical properties, the model stays schema-flexible instead of locked to a fixed set of columns.
Relationships work the same way. They can carry their own properties, a KNOWS edge might carry since: 2020, and they can point one way or both, unlike a foreign key, which always points in exactly one direction.
Querying with Cypher
CREATE (r:Person {name: "Rohit Sharma", born: 1987})one labeled node, two properties, no INSERT and no schema migration
Neo4j's query language, Cypher, is declarative and visual: a query is written to look like the pattern being searched for. CREATE (r:Person {name: "Rohit Sharma", born: 1987}) creates one labeled node with two properties in a single statement, no INSERT, no separate schema migration first.
Retrieval reads the same way. MATCH (p:Person)-[:CAPTAIN_OF]->(t:Team) RETURN p, t matches the shape "a Person, connected to a Team by a CAPTAIN_OF edge," wherever that shape occurs in the graph. Ask most data engineers what the hardest part of SQL is and the answer is usually joins and nested subqueries, Cypher has neither.
Multi-hop questions, the kind that need several joins in SQL, become one pattern. Match a Person connected to a Team, then match back out to a second Person connected to that same Team, and the two people the query lands on are teammates by construction, with no separate teammates table required anywhere.
Building one from raw text
Everything so far assumes the graph already exists. Building one from scratch used to mean a domain expert manually reading source material and labeling entities and relationships by hand, or a rule-based extractor built to catch a narrow set of patterns, brittle, and useless the moment the text didn't match the rules it was written for.
Statistical NLP and machine learning closed part of that gap, but only part: most models were trained overwhelmingly on English, and they still stumbled on the nuance, context, and ambiguity plain language is full of.
Modern LLMs close most of what's left. Pointed at raw, unstructured text, in any of the languages they were trained on, they can read for context and return the entities and relationships in it directly, which is what makes this practical at the scale a real knowledge graph needs. To see it without writing any code, Neo4j's free LLM Knowledge Graph Builder does exactly this through a web page: paste in text, get a graph back.
Try it: extract the triple
Read a sentence the way a knowledge-graph pipeline would, then pick the node-edge-node triple that actually represents it.
Read the sentence, then pick the triple that correctly represents it as a graph edge.
"Marie Curie won the Nobel Prize in Physics in 1903."
Waiting for your guess.