GFQL Cypher Filter + PageRank Benchmark#

GFQL mascot

Run Cypher queries and graph analytics directly on Python dataframes, without a database. This benchmark compares Graphistry’s local Cypher on CPU and GPU with Neo4j + GDS for the same three-stage pipeline.

Neo4j + GDS

GFQL Cypher (CPU)

GFQL Cypher (GPU)

GFQL GPU vs CPU

Twitter (81,306 nodes / 2.4M edges)

11.72 s

1.58 s

0.24 s

6.7x

GPlus (107,614 nodes / 30M edges)

354.47 s

32.10 s

2.42 s

13.3x

Each time covers the full search → PageRank → search pipeline after warm-up. GFQL reuses data already loaded in Python. Neo4j includes server calls and rebuilds the in-memory graph used by Graph Data Science (GDS) for each timed iteration. The table therefore shows direct pipeline times, not a GFQL-to-Neo4j speedup ratio.

For the same GFQL query, the GPU path is 6.7x faster on Twitter and 13.3x faster on the 30M-edge GPlus graph.

The pipeline#

One g.gfql(...) call searches the graph, calculates PageRank, and searches the result:

# pip install graphistry
result = g.gfql("""
    GRAPH g1 = GRAPH {
      MATCH (n)-[e]-(m)
      WHERE n.degree >= $degree_cutoff
    }
    GRAPH g2 = GRAPH {
      USE g1
      CALL graphistry.cugraph.pagerank.write()
    }
    GRAPH {
      USE g2
      MATCH (n)-[e]-(m)
      WHERE n.pagerank >= $pagerank_cutoff
    }
""",
    params={
        "degree_cutoff": degree_cutoff,
        "pagerank_cutoff": pagerank_cutoff,
    },
    engine="cudf",  # or "pandas" with igraph backend
)
  • GRAPH g1: find high-degree nodes and their neighbors

  • GRAPH g2: enrich g1 with PageRank scores (igraph on CPU, cugraph on GPU)

  • Final GRAPH: keep high-PageRank nodes and their neighbors

Choose a CPU or GPU backend without changing the query:

  • CPU: engine="pandas", backend="igraph"

  • GPU: engine="cudf", backend="cugraph"

The Neo4j version requires Cypher, a separate in-memory graph for GDS, and several writes. See Neo4j + GDS analog below.

Twitter (2.4M edges): reported pipeline timings#

Twitter warm pipeline time: Neo4j + GDS 11.72s, GFQL Cypher CPU 1.58s, GFQL Cypher GPU 0.24s
  • Neo4j + GDS: 11.72 s

  • GFQL Cypher on CPU (pandas + igraph): 1.58 s

  • GFQL Cypher on GPU (cuDF + cuGraph): 0.24 s — 6.7x faster than the GFQL CPU path

GPlus (30M edges): larger graph#

GPlus warm pipeline time: Neo4j + GDS 354.47s, GFQL Cypher CPU 32.10s, GFQL Cypher GPU 2.42s
  • Neo4j + GDS: 354.47 s

  • GFQL Cypher on CPU (pandas + igraph): 32.10 s

  • GFQL Cypher on GPU (cuDF + cuGraph): 2.42 s — 13.3x faster than the CPU path

GPlus is 12x the edges of the Twitter graph, and the GPU pipeline still answers in seconds.

What this shows#

GFQL runs the same query on pandas + igraph or cuDF + cuGraph. The GPU path was faster on both graphs. GFQL also keeps dataframe processing, graph search, and analytics in one Python process.

Neo4j + GDS analog#

The Neo4j equivalent of the same pipeline:

-- 1. Mark seed nodes by degree
MATCH (n:Node)
SET n.seed = n.degree >= $cutoff;

-- 2. Expand one hop from seeds
UNWIND $seed_ids AS sid
MATCH (s:Node) WHERE id(s) = sid
MATCH (s)-[r:LINK]-(target:Node)
SET target.in_subgraph = true, r.in_subgraph = true;

-- 3. Project subgraph and run PageRank
CALL gds.graph.project.cypher(
  'subgraph',
  'MATCH (n:Node) WHERE n.in_subgraph RETURN id(n) AS id',
  'MATCH (a)-[r:LINK]->(b) WHERE r.in_subgraph
   RETURN id(a) AS source, id(b) AS target
   UNION ALL
   MATCH (a)-[r:LINK]->(b) WHERE r.in_subgraph
   RETURN id(b) AS source, id(a) AS target'
);
CALL gds.pageRank.write('subgraph', {writeProperty: 'pagerank'});

-- 4. Keep high-PageRank core + one hop
MATCH (n:Node) WHERE n.pagerank >= $cutoff
SET n.core = true;
UNWIND $core_ids AS cid
MATCH (c:Node) WHERE id(c) = cid
MATCH (c)-[r:LINK]-(target:Node)
SET target.final = true, r.final = true;

Why the GFQL pipeline is shorter#

The Neo4j version is longer because its stages write flags to database records and create a separate GDS graph. GFQL passes a graph directly from one stage to the next.

Graphs as values. Each GRAPH { } block receives a graph, changes it, and passes a graph to the next block. This removes the property flags, separate GDS projections, and batched writes used in the Neo4j example.

One query, multiple engines. GFQL compiles Cypher to dataframe operations. Set engine="pandas" for CPU execution or engine="cudf" for GPU execution. See Cypher Syntax In GFQL for supported Cypher features and Overview of GFQL for the GFQL design.

Columnar data in Python. Intermediate graphs stay in Arrow, pandas, or cuDF memory. ETL, search, and analytics can remain in the same Python pipeline.

Consistent results. GFQL either returns the same result on an engine or rejects the query before execution. It does not silently change engines. See Choosing a GFQL Engine: pandas, Polars, cuDF, Polars-GPU for the parity and validation rules.

This page is one workload (a filter → PageRank → filter pipeline) against one external baseline (Neo4j + GDS). For the full four-engine picture — when Polars beats pandas on CPU, when the GPU pulls ahead, and how to choose — see Choosing a GFQL Engine: pandas, Polars, cuDF, Polars-GPU. For seeded lookups, see Seeded Traversal Indexes (CSR Adjacency).

For more on GFQL:

Benchmark environment and provenance#

Every figure is printed from docs/source/_data/gfql_benchmarks.json (pyg-bench).

Measurement

Measured:

2026-07-28

Host:

dgx-spark (NVIDIA GB10, driver 580.126.09), 20 CPU

Repetitions:

graph loaded once, then 2 warmups + 5 timed runs per arm on the resident graph; median

Runtime:

graphistry/test-rapids-official:26.02-gfql-polars with python-igraph 1.0.0; cuDF 26.2.1, cuGraph 26.2.0, pandas 2.3.3; Neo4j 2026.02.2 + graph-data-science in Docker on the same host

Dataset:

SNAP twitter_combined (81,306 nodes / 2,420,766 edges) and gplus_combined (107,614 nodes / 30,494,866 edges); sha256 of each source file is recorded in the arm artifacts

PyGraphistry commit:

49db91cc

Benchmark commit:

85c92022 plus benchmarks/filter_pagerank as added in this commit

Raw artifacts:

results/filter-pagerank-20260728

Result validation:

Every arm records the node id set its pipeline selected, captured outside the timed region. Comparability is the Jaccard index of those sets against a 0.95 threshold declared before the run: Twitter CPU/GPU 0.991, CPU/Neo4j 0.974, GPU/Neo4j 0.972; GPlus CPU/GPU 0.951.

Competitor version:

neo4j:2026.02.2 with the graph-data-science plugin

Measurement

Measured:

2026-08-30

Host:

dgx-spark (NVIDIA GB10, driver 580.173.02), 20 CPU

Repetitions:

12 position-balanced slots, six per arm; each slot 2 warmups + 11 timed runs; median of slot medians

Runtime:

graphistry/test-rapids-official:26.02-gfql-polars; GFQL Python 3.13.12, pandas 2.3.3, python-igraph 1.0.0; Neo4j 2026.02.2 + graph-data-science with Python client 3.12.3 on the same host

Dataset:

SNAP gplus_combined (107,614 nodes / 30,494,866 edges; 398,930,514 bytes; sha256 492b7a63ec7816cac6aa0466be8521a0528cb1df2ca7f0b554ad746adc6bbbee)

PyGraphistry commit:

76dc3f30242c2b956b5a7e133a32fb80495ef9e3

Benchmark commit:

736f1b8fc219a23f38d5e32e5e291932ec2aaf2f

Raw artifacts:

results/gplus-locked-baseline-20260830

Result validation:

All 12 selected-node sets had Jaccard 1.0 against the reference (gate 0.95); every slot passed typed result, exact run-contract, and load/self-spike validation; the committed aggregate exactly recomputes from all slots.

Competitor version:

neo4j:2026.02.2 with the graph-data-science plugin

About these measurements

  • In the Neo4j arm each filter stage writes a marker property onto every matched node and relationship, 32 seeds per batch, and GDS scores an explicitly symmetrised projection; the GDS in-memory projection is rebuilt per timed iteration while the GFQL arms retain their resident frames. The GFQL arms materialise a subgraph and write nothing. The figure is a pipeline time, not an engine-primitive time.

  • The Neo4j arm writes marker properties during both filter stages and rebuilds its GDS projection per timed iteration; the paired GFQL arm retains resident frames and writes nothing. Twelve position-balanced slots (six per arm) produced exact selected-node parity, but the measurement profiles differ, so this is a direct pipeline time and no GFQL-vs-Neo4j ratio is published.