A graph database stores nodes (entities) and edges (relationships) and is optimised for traversal queries — "find all nodes reachable from this one within 3 hops." The two main flavours: property graphs (Neo4j-style) and triple stores (RDF-style).
Most teams considering a graph database don't actually need one — Postgres with a graph schema handles 90% of cases. This page is the working set for when you genuinely do, and what to expect.
Each node has:
:Person, :Company).Each edge has:
WORKS_AT).(p:Person {name:'Dario'}) -[:WORKS_AT {role:'CEO', since:2021}]-> (c:Company {name:'Anthropic'})
This is what Neo4j, JanusGraph, TigerGraph, and most modern graph DBs use.
Neo4j's Cypher has become the de-facto standard for property-graph queries. Pattern-matching syntax:
MATCH (p:Person)-[:WORKS_AT]->(c:Company {name:'Anthropic'})
RETURN p.name, p.role
Powerful for graph patterns:
// All people who work at companies that compete with companies in SF
MATCH (p:Person)-[:WORKS_AT]->(c1:Company)-[:COMPETES_WITH]->(c2:Company {city:'San Francisco'})
RETURN p.name, c1.name, c2.name
Variable-length paths:
// Find shortest path between two people through any relationship
MATCH path = shortestPath((a:Person {name:'Alice'})-[*..6]-(b:Person {name:'Bob'}))
RETURN path
This is where graph DBs shine. The equivalent SQL with recursive CTEs is verbose and slower for deep traversals.
GQL (Graph Query Language) standardised in 2024 is the ISO standard inspired heavily by Cypher. Most graph DBs are converging on GQL or maintaining Cypher compatibility.
The useful test is not "is my data connected?" — all data is connected. It's: do my most valuable queries follow chains of relationships whose depth or shape I don't know in advance? When the answer is yes, a graph database stops being a luxury and starts paying for itself. The domains below are the ones where that reliably happens. For each, note the killer query — the one that is natural in Cypher and miserable in SQL.
Money laundering and organised fraud hide in paths, not rows: shell accounts, shared devices, shared addresses, and chains of transfers that connect a new account back to a known-bad one.
MATCH path = (a:Account {flagged:true})
-[:TRANSFER|SHARES_DEVICE|SHARES_ADDRESS*1..6]-
(b:Account {id:$newAccount})
RETURN path LIMIT 1
"Customers like you also bought…" is a multi-hop traversal over a bipartite user–item graph, computed live from a seed node rather than from a pre-baked table.
MATCH (me:User {id:$u})-[:BOUGHT]->(:Product)<-[:BOUGHT]-(peer:User)-[:BOUGHT]->(rec:Product)
WHERE NOT (me)-[:BOUGHT]->(rec)
RETURN rec.name, count(*) AS score
ORDER BY score DESC LIMIT 10
Authorization is a reachability question: can this principal reach this resource through any chain of group memberships, roles, and grants? This is exactly the model behind relationship-based systems like Google Zanzibar.
MATCH (u:User {id:$u})-[:MEMBER_OF|HAS_ROLE|GRANTS*1..8]->(r:Resource {id:$y})
RETURN count(*) > 0 AS allowed
Telecom networks, microservice meshes, IT asset inventories, and software dependency trees all ask the same thing when something breaks.
MATCH (svc:Service {id:$s})<-[:DEPENDS_ON*1..]-(affected)
RETURN DISTINCT affected.name
Manufacturing data is inherently multi-tier: components roll up into sub-assemblies into finished goods, across many supplier levels.
Records describing the same real-world person, company, or device arrive from many systems and must be linked into one view.
Influence, communities, and "degrees of separation" are graph-algorithmic by definition.
Entities plus typed relationships power semantic search and retrieval-augmented generation for LLMs, where traversing relationships expands the relevant context.
Shortest/cheapest path and max-flow are classic graph problems, and graph DBs handle them well.
The common thread: every case above is variable-depth reachability, pathfinding, or a graph algorithm — not a fixed-shape report. If your headline queries fit that description, a graph database earns its place.
"We have relationships" is not a reason — every relational schema has foreign keys. The honest question is whether traversal is your bottleneck. Where it isn't, another store is simpler, cheaper, and faster. Each anti-pattern below names the tool that actually fits.
ltree. A graph DBMS is overkill when depth is small and known.Rule of thumb: if you can't name a query that traverses a variable or unknown number of hops, you probably don't need a graph database.
Score your project. The more rows you land on the left, the more a graph database earns its keep.
| Lean graph database | Lean another store |
|---|---|
| Headline queries traverse 3+ hops, often variable-length | Almost every query is 1–2 hops |
| You need pathfinding, shortest-path, or cycle detection | You mostly need GROUP BY / SUM over big tables |
| Relationships carry properties you actually query | Relationships are just foreign keys you join on |
| You'll run graph algorithms (PageRank, community, centrality) | You need full-text search or time-series rollups |
| The shape of connections is irregular and evolving | The schema is tabular and stable |
| Connected-data queries are your competitive core | Connected data is incidental to a mostly-CRUD app |
Verdict: 4+ ticks on the left and a graph database likely pays off. Mostly right-hand ticks — start with Postgres (a nodes/edges schema, or the Apache AGE extension for Cypher syntax) and migrate only if a concrete query later forces the move.
The dominant graph database. Mature; rich ecosystem; excellent Cypher implementation; Graph Data Science library for algorithms.
For most graph DB projects: Neo4j is the safe default.
High-performance distributed graph DB. Strong on analytical workloads at scale. GSQL query language. Commercial.
For very large graphs (hundreds of millions of nodes) and analytical workloads, TigerGraph competes well.
Open-source distributed graph; built on Cassandra / HBase / ScyllaDB for storage. Gremlin query language (TinkerPop standard).
Strengths: very large scale; open source; flexible storage backend. Weaknesses: complex to operate; Gremlin is harder than Cypher; less polished than Neo4j.
Newer; Cypher-compatible; in-memory; fast. Open source core; commercial features.
Good for streaming graph analytics; less mature than Neo4j for general use.
Multi-model: graph + document + key-value. Useful when you genuinely need multiple shapes; less specialised than dedicated graph DBs.
These let you avoid running a separate graph DB. For mid-scale graph workloads inside a Postgres-shop, they're worth trying.
A different model: triples (subject-predicate-object) with formal semantics, ontologies (OWL), and SPARQL queries.
Used in:
Tools: Apache Jena Fuseki, Stardog, GraphDB, Virtuoso.
For most modern industrial applications, property graphs win on usability. RDF is more powerful for formal reasoning but the property-graph model is simpler.
A few patterns that age well:
Edges are directed; queries can traverse either direction. Cypher's --> vs <-- vs --. For most relationships, directionality matches the natural reading: (employee)-[:WORKS_AT]->(company).
Properties on edges represent relationship attributes — since, role, weight, confidence. Use them; don't promote every edge property to a separate node.
A node with millions of edges (a celebrity in a social graph; "USA" in a "located in" graph) creates query hotspots. Some traversals balloon at super-nodes.
Mitigations: split logically (separate edge types for different relations); index the edge type to filter early; avoid traversing through super-nodes by structure.
Multiple labels per node give flexibility:
(p:Person:Engineer {name:'Dario'})
Queries can match :Person, :Engineer, or both. More expressive than a single label.
EXPLAIN.For a new project considering a graph DB:
Most teams should start with Postgres + graph schema. Migrate to a dedicated graph DB if and when specific limits force it.
What is a graph database? A database that stores data as nodes (entities) and edges (relationships) and is optimized for traversal queries — following relationships across many hops — rather than the row/table joins of a relational database.
Graph database vs. relational database — what's the difference? Relational databases excel at structured, tabular data and set-based queries; graph databases excel at deeply connected data and multi-hop traversals where SQL joins become slow and verbose. See Knowledge Graph vs. Relational Database.
When should I use a graph database? When you run frequent deep traversals (5+ hops), need built-in graph algorithms (PageRank, community detection, shortest path), or model richly connected domains like fraud rings, recommendations, and social networks. For shallow 1–2 hop queries, Postgres is usually enough.
What are some real-world graph database use cases? Fraud-ring and anti-money-laundering detection, recommendation engines, identity and permission (reachability) checks, network and dependency/impact analysis, supply-chain and bill-of-materials roll-ups, master-data and entity resolution, social-network analysis, and knowledge graphs / GraphRAG. The common thread is queries that follow chains of relationships of variable or unknown depth.
When is a graph database overkill — when should I NOT use one? When your queries are shallow (1–2 hops), aggregation-heavy (use a columnar warehouse), document- or key-value-shaped, time-series, or search-centric — or when relationships are merely foreign keys you join on. If you can't name a variable-depth traversal query, skip the graph database.
Can I run graph queries in Postgres instead?
Yes. A nodes/edges schema with recursive CTEs covers shallow-to-moderate graphs, and the Apache AGE extension adds Cypher syntax on top of Postgres. Most teams should start here and only migrate to a dedicated graph DB when a real query or scale limit forces it.
What is the most popular graph database? Neo4j is the dominant property-graph database and the safe default for most projects; TigerGraph, JanusGraph, Memgraph, and ArangoDB serve more specialized needs.
What query language do graph databases use? Cypher (Neo4j and compatibles) is the de-facto standard; the ISO GQL standard (2024) is heavily inspired by it. JanusGraph uses Gremlin; TigerGraph uses GSQL.