docs(gfql): publish final performance results and simplify engine guides - #2017
Conversation
41bf7fe to
350ef4b
Compare
350ef4b to
cab184f
Compare
|
READY (cab184f) but LAST per the landing order: rebased on master; 59 green + 10 skipped + RTD; re-vendor from the master re-measure follows before the final merge. |
|
Re-vendored from the master re-measure: docs/source/_data/gfql_benchmarks.json now equals pyg-bench published/docs-numbers.json at pyg-bench main af1513f (PR graphistry/pyg-bench#249: SNB SF0.1/SF1 arms and q1–q9 20k/100k boards at pygraphistry master 5a6586f, sentinel baseline at 86de0f5). 166 of 262 cells changed, all within noise of the 4 September board except the seeded SNB shapes, which are the release wins (pandas message-content 10.06→1.66 ms, seed-lookup 30.12→3.91 ms; polars seed-lookup 21.82→10.19 ms). Losses stay on the pages through the same cells (100k q8 Kuzu 9.7 ms vs polars 14.0 ms; polars tag-cooccurrence 30.2→31.7 ms, inside its 17% run-to-run spread). The SNB pre-landing disclosure is gone (both scales are the release measurement); the GraphFrames note now says #2024 has landed since and the ladder was not re-run. docs/test_bench_numbers.py 37 pass. Stays open until the rest of the queue lands (merges last). |
|
CI receipt at 090ad9e: 59 check-runs success, 10 skipped by the path filter (docs-only change). docs/test_bench_numbers.py 37 pass locally; the vendored JSON equals pyg-bench main af1513f. |
090ad9e to
60d39f0
Compare
|
Rebased onto master 1a41079 (60d39f0, docs only, clean). Expect the docs-numbers contract test to be RED on this head by design: the vendored boards were measured at 5a6586f and the landed 0.60 stack put them 19 compute commits behind master (policy allows 12). Rather than waive, the SNB arms, q1–q9 boards and sentinel are being re-measured at master 1a41079 on dgx now; this PR gets re-vendored from that run, with any cell that moved called out here. The three 0.59.0-era runs (GraphFrames ladder ×2, filter-pagerank) get explicit drift waivers with the reason stated in the pyg-bench PR. Review the prose and structure via the RTD preview meanwhile: https://pygraphistry--2017.org.readthedocs.build/en/2017/gfql/index.html |
|
CI on 60d39f0: 59 green, 11 filter-skipped. Correction to my note above: CI does not run docs/test_bench_numbers.py on this PR (it is among the skips), so CI green here is not evidence for the numbers. Locally that test fails on drift (19 compute commits past the policy's 12) until the re-vendor from the master 1a41079 re-measure lands. |
|
Editorial pass 1 (owner items 1–5, 7–12 on engines; 6 started) at 3dbf8fc. Preview rebuilds from this head in a few minutes:
|
|
Jargon pass (item 6) at 777475d across performance, index_adjacency, indexing, benchmark_filter_pagerank, benchmark_graphframes, overview: 'seeded' → 'from known nodes' except the two defining first uses; 'receipt' / 'committed artifact' → 'run record'; 'parity' → 'identical results'; 'point lookup' → 'single-node lookup by id'; 'lane' → 'path'. No number changed. Items 1–12 from your list are now all applied; add more per page and I queue them in plan.md before editing. |
|
Re-vendored at d660e33 from the master 1a41079 re-measure (pyg-bench #251, publishing now). docs/test_bench_numbers.py passes locally including the drift policy (37 passed). Cells that moved more than 15% vs the 5a6586f vendoring: graphbench.20k.q4.polars_gpu 16.41 → 4.53 ms (the GPU tiny-point outlier did not reproduce), graphbench.20k.q8.polars_gpu 2.17 → 2.53, snb.sf01.message_creator.gfql_pandas_idx 2.87 → 2.33; everything else within run-to-run spread. Verdict changes on the boards: 100k q5 polars vs Kuzu is now a tie (0.97×, was a weak 1.12× win); 20k q8 is a tie (1.10×); 100k q8 stays the one loss. Board prose on performance.rst updated to the new verdicts; the tallies are roles and follow the cells automatically. |
d660e33 to
762b011
Compare
a489e9a to
b3dd07c
Compare
…chmark pages - gfql/index: Start Here (about, overview, quick, cypher) + Guides hubs gfql/perf/index and gfql/reference/index + Developer Resources; no page moved, URLs unchanged; loading_graph_data joins the reference hub - benchmark_filter_pagerank, benchmark_graphframes: lede states result and engine to use; one "Method and limits" section; numbers unchanged and re-verified against bench cells / results.json; Cypher comments use // - CHANGELOG: [Development] / Docs Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
…provenance block - overview: positioning paragraph (only open-source in-process Cypher on dataframes; columnar/vectorized framing); remove empty hidden toctree that rendered the page as a folder; trim marketing wording - benchmark_filter_pagerank: retitle as a case study, list it under Start Here; lede says what is compared and the outcome; drop the above-the-fold caveat paragraph; add a pipeline lead sentence - benchmark_graphframes: lede says what is compared and the 7-of-8 outcome; add per-task bar charts rendered from results.json by gfql_bench_charts.py (byte-reproducible, covered by the chart sync test) - _ext/gfql_bench: bench-provenance accepts several runs and renders one Measurement block (reader-facing fields only; caveats folded in via :disclosures:) - vendor pyg-bench published numbers (runtime strings now name the RAPIDS base image; pyg-bench #220); numbers unchanged Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
…e 6/8); wrap chart strings Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
…engines table, tested Cypher twins - index: Start Here = about, overview, speedup case study; quick + cypher head the Language Reference hub (spec linked) - benchmark_filter_pagerank: "Speedup Case Study" title - indexing / index_adjacency: plain-language titles and cross-refs - overview: one "Why GFQL?" list (was Why + Key Features), engine-design and launch blog links - engines: opening note boxes folded into prose; "coming from" table with concrete change + measurement pointer; Memgraph row; PuppyGraph removed; section "Parity and fallback rules" - about: examples 3-7 gain tested Cypher twins; example 4 pattern now returns rows on the sample graph; sample-graph block is executable in the doc lane - slop words removed across gfql pages (leverage/seamless/honest/powerful/ critical gap); no numbers added Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
…NB matrix; cell-derived tallies; drift enforcement - vendor pyg-bench published/docs-numbers.json @ main 7426456 (relane board with Kuzu/Memgraph/Neo4j, snb_aligned SF0.1/SF1, drift waivers) - _ext/gfql_bench_data: :bench-tally: (strict "N of M" from cells, registers refs) and max_compute_commit_drift enforcement via git rev-list on graphistry/compute, honoring policy.drift_waivers; None on shallow clones - performance.rst: no release-pinned headings; legacy untraceable literals removed; SNB tables with losses stated; five-engine q1-q9 boards; single Measurement block with caveats; q8 rendered as a result (cold binding) - overview: where GFQL wins and where databases win, with links - tests: tally, drift (fail / waived / unknown), vendored runs within policy Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
…s diagnostic - vendor pyg-bench published numbers with results/filter-pagerank-059-20260904 (GFQL CPU/GPU re-measured at 3fb216d under the pinned PageRank contract) - lede states CPU-vs-Neo4j on both graphs and the Twitter GPU ratio; the GPlus GPU time renders via :bench-diag: with the selection caveat, and no GPlus GPU-vs-CPU ratio is claimed; alt texts follow the cells - gplus_pipeline chart drops the withdrawn ratio annotation Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
Vendor the pyg-bench publication (graphbench-q1q9-{20k,100k}-20260904:
pandas, polars, polars-gpu at 3fb216d; Memgraph/Neo4j lanes of 2026-08-12).
performance.rst gains the polars-gpu column, tallies for GPU-vs-CPU and
GFQL-vs-each-database, and names every loss (Kuzu q4/q8, Memgraph q3/q5/q6/q7,
Neo4j q5, GPU q8). Provenance block retargeted to the new runs.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
b3dd07c to
fb51a52
Compare
Rebased onto the shipped master, and every number the re-vendor needs now existsRebased onto This PR's own gate proves the re-vendor is due
And they carry pre-release numbers. The currently vendored Measured on the shipped tree, ready to vendorAll receipted under the DGX idle gate and host perf lock, harness H684 SNB, polars — SF0.1 seed-lookup 0.222, message-replies 1.355, recent-replies 5.466, new-topics 42.463, message-content 0.165, message-creator 0.224. SF1 seed-lookup 0.281, new-topics 168.253, message-content 0.164, message-creator 0.232. SNB, pandas — SF0.1 seed-lookup 0.913, message-replies 11.810, recent-replies 27.222, new-topics 72.171, message-content 0.424, message-creator 0.621. SF1 seed-lookup 1.147, new-topics 576.991. GraphBench — 20k 8 win / 1 tie / 0 loss; 100k 7 win / 1 tie / 1 loss (q8, Two things the copy must say
What is leftDeclaring these runs in a pyg-bench spec entry, regenerating |
…tree The vendored numbers were pre-release and four runs were 52-61 compute commits past their measuring commit, against a policy limit of 12 -- this PR's own drift gate was failing on them. seed-lookup published 1.018 ms at SF0.1 where the shipped tree measures 0.218, and 1.363 at SF1 where it measures 0.281. Re-vendored from pyg-bench with SNB measured on pygraphistry f283a30 (master with #2084, #2086, #2087, #2088, #2090) and GraphBench on 24c0b1e, both under the DGX idle gate and host perf lock with canonical rows asserted identical across engines, scales and repetitions. docs/test_bench_numbers.py now passes: 37 passed, 1 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
The previous vendor could not be regenerated. Its own generated_by named pyg-bench `baf58bd8`, which is not an ancestor of main -- it came from the lineage of closed PR #271 -- and the `results/graphbench-*-24c0b1eba` directories it depended on existed on no surviving branch. The drift gate could not see this: that board sat at drift 4 against a limit of 12. It also published GraphBench from the four-PR tree while publishing SNB from the shipped one. Re-vendored from pyg-bench main `e7c49f44`, where GraphBench is now published from the shipped release tree (pyg-bench #277). All four measured runs sit at compute-drift 0; the remaining five carry standing waivers. Every vendored run's artifact directory is present on pyg-bench main, so these numbers are reproducible by anyone. docs/test_bench_numbers.py: 37 passed, 1 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…carries, and pin it test-docs was already failing at the previous head: six `bench-provenance` arguments named run ids absent from the vendored artifact. `bench-provenance` takes run ids as literal text, so a republished board silently orphans them and only a full docs build notices. Four of the six were stale before today's republication -- performance.rst still named snb-*-master-f7a7253bc-20260913 and graphbench-*-d20c6ae1a-20260914, none of which the previous vendor contained either. The other two moved when GraphBench was republished from the shipped tree. Adds test_every_bench_provenance_names_a_published_run so this fails in seconds rather than in a ten-minute Sphinx build. Verified by reinstating a stale id: the pin fails, and passes again when restored. docs/test_bench_numbers.py: 38 passed, 1 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
pyg-bench #280 re-measured the ladder's LiveJournal and Orkut GFQL arms at pygraphistry 62df29a, each source in its own published mode, all sixteen cells within +-6.6% of the board they replace. Re-vendored from pyg-bench main 01353d84. Two contract tests flagged the consequences and both are fixed here: * bench-provenance still named graphframes-ladder-20260904, a run id that no longer exists. Caught by test_every_bench_provenance_names_a_published_run, the pin added when six such references went stale at once on the GraphBench republish -- it has now paid for itself. * livejournal_tasks.svg and orkut_tasks.svg no longer matched the published numbers; regenerated from the artifact. Vendored drift is within policy everywhere: the ladder at 0, GraphBench and SNB at 4, and the five standing waivers unchanged. docs/test_bench_numbers.py: 38 passed, 1 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
The PR had gone CONFLICTING against master (15 commits ahead), which fires ZERO CI runs and looks exactly like a stuck webhook: the head was pushed, the PR pointed at it, and no workflow existed for it. One real conflict, in docs/source/gfql/cypher.rst, and it is not this PR's subject. Master rewrote the searchAny prose substantively (integer/float/datetime matching, inspector rendering, temporal_tz and the cuDF zone decline). This branch had only an editorial change to the same paragraph: dropping the word "honestly" from two phrases. Resolved by taking master's content and RE-APPLYING the branch's edit on top, so both intents survive rather than one silently winning. Vendored artifact and contract tests unaffected: 9 runs including the republished ladder, docs/test_bench_numbers.py 38 passed / 1 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
The engines page opens "GFQL runs the same query on four execution engines" and is titled "pandas, Polars, cuDF, Polars-GPU", but its switch-engines example listed only auto, polars, cudf and polars-gpu. A reader pinning pandas explicitly -- to stop auto following Polars frames, say -- had no line to copy. Adds it there, and to the "Selecting an Engine Explicitly" block in overview.rst, which had the same gap; its lead sentence is widened to match, since it had scoped the block to "a CPU columnar speedup or ... a specific GPU engine". Checked the other pages rather than assuming: notebooks/gpu.rst and the about.rst block are titled for GPU engines specifically, so omitting pandas there is correct and they are left alone. Wording follows the engines table on this same page: "CPU execution with pandas frames". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
The "Coming from" table on the engines page carried a LadybugDB row. Competitor discussion in the docs should be Kuzu, Memgraph, Neo4j and the non-database tools only. Removing it orphans nothing: the :ref:`gfql-larger-than-memory` target it pointed at is still linked from the same page's capacity row and from gfql/perf/index.rst, both checked before the removal. docs/ now contains zero LadybugDB references. Three non-docs occurrences are deliberately left alone: the CHANGELOG entries, which RECORD the withdrawal of unverifiable LadybugDB figures and the later real measurement -- rewriting them would erase a correction that was deliberately called out rather than buried; the two benchmarks/gfql/bench_ladybug*.py runners, which are code rather than documentation; and a comment in graphistry/compute/ast.py, which is compute-path and would add drift to a docs PR that must land before the boards go stale. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Owner editorial round 1 on performance.rst: retitle to "GFQL outperforming traditional graph databases", lead rewritten, "q1-q9 board" -> "graph-benchmark", "SNB-derived lookups and small-result queries" -> "SNB Interactive", and the protocol/methodology prose (warmup counts, slot balancing, index build, adapter exclusions) dropped from the page. Numbers still render from the published artifact via the bench roles, and the provenance block keeps date and hardware. Two items asserted numbers the page's own data refutes, so both were reworded rather than published as given: - "< 5ms" for CPU queries: the boards show 2.15-9.68 ms at 20k (6/9 under 5) and 5.62-37.95 ms at 100k (0/9 under 5). Reworded to "answers these queries in milliseconds", which keeps the point that GPU buys nothing at this size. - A blanket SNB win over all three competitors: true everywhere except recent replies at SF0.1, where GFQL 5.387 ms vs Neo4j 5.425 ms is parity. Stated as the one exception. bench-provenance gains an opt-in :fields: option so this page can show date and hardware alone. Four pages use the directive, so narrowing the shared field list would have reshaped three pages that were not in scope. An unknown field name raises rather than dropping a row. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
benchmark_graphframes: lead with the result instead of the method. The old "Where it stands" paragraph ran twelve lines before reaching a takeaway; it is replaced by one claim (GFQL queries billion-edge graphs on one machine, with Orkut's 117M edges filtered and expanded in under a second) and a four-bullet win/loss list. Both sides are still printed: GPU PageRank and CPU filter/1-hop are wins, 2-hop and CPU PageRank are losses. Binding/protocol prose, the GPU streaming aside, and three micro-details are dropped, and exact counts become scale figures (3,997,962 -> 4.0M). Charts now order GPU, CPU, GraphFrames. benchmark_filter_pagerank: the opening block is replaced by the pipeline shape and the magnitudes, minutes to seconds and seconds to subsecond. The GPlus GPU figure stays out of the summary: its cuGraph PageRank selects a different node set than igraph, so the cell is board_quotable=false and cannot become a headline ratio. Both provenance blocks narrow to date and hardware via :fields:. The ladder examples use the native op-list surface rather than Cypher. Checked whether that forces a re-run: a Cypher compile is 0.94-2.61 ms, which is 0.002% to 10.6% of these cells, and adding 2.6 ms to every GFQL cell flips no verdict (the tightest, Orkut filter at 57.7 vs 66.1 ms, keeps a margin three times the compile cost). The numbers stand as published. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Applies the same :fields: narrowing to the two other pages that carry a provenance block, so all four are consistent rather than three showing a short block and one showing the full field list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Removes :disclosures: from all four provenance blocks, so the pages no longer recite warmup counts, execution counts, slot balancing, adapter exclusions or the LDBC-driver note. The provenance block is now date and hardware only, as asked. The disclosures stay attached to the cells in the published artifact; they are simply no longer printed on the page. Also drops "massive speedups" from about and overview, and updates overview's SNB sentence to match the wording performance.rst now uses, including the SF0.1 recent-replies tie. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Removing :disclosures: broke the build. gfql_bench_data.py raises when a page prints a number whose cell carries a disclosure and the page declares none: "prints a number that carries a disclosure but has no bench-disclosures block". 114 published cells carry disclosures, and three of these pages print them (graphframes 41, performance 4, filter_pagerank 3), so all three failed and the docs did not publish. The protocol text still needs to go. It cannot go from the page alone: the text lives in the disclosure strings of the published artifact, so the fix is in pyg-bench, separating true validity caveats from protocol description, then a re-vendor. Restoring the declaration here only gets the docs building again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
The install line named only pandas and cudf; Polars is a first-class engine now, so it lists all three and points at the engines page for the choice. The introduction and Key Benefits are rewritten as positioning: GFQL is built for developers and AI coding agents, and for typical graph tasks is faster, easier, and safer than a graph database, without the infrastructure sold alongside one. Each claim is tied to a mechanism or a measured page rather than asserted: "faster" links to performance (Kuzu, Memgraph, Neo4j, Spark GraphFrames), "safer" names pre-execution validation and structured error codes and links to the validation pages, "easier" is pip install and nothing to deploy. No new performance figure is written on this page; it carries no bench roles, so any number here would be unverifiable against the artifact. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…eads "all" overview: the CPU-or-GPU sentence now carries the positioning (faster than the measured graph databases, linked to performance; no database to stand up). Key Concepts states the data sources a graph loads from (CSV, Parquet, JSON, SQL such as Postgres, graph databases such as Neo4j) and links loading_graph_data. Splunk was requested but is not a loader anywhere in the library or docs, so it is not listed. The cypher() warning now says plainly what those calls are: PyGraphistry's BOLT-driver bindings to Neo4j, Memgraph, or Neptune, not GFQL. The inline dump of twelve NetworkX .write() enrichers is replaced by one sentence and a link to builtin_calls, where the list is maintained. bench-tally: a clean sweep renders as "all" (and "none"), not "9 of 9". Role level, so engines and performance both pick it up; every tally site was read with the substitution and stays grammatical. Pinned in test_bench_numbers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Quick start now opens with Cypher: CREATE GFQL INDEX FOR edge_out_adj, a seeded MATCH, and SHOW GFQL INDEXES, with the native-chain form second. The example frames are defined inline so the block is self-contained, and the two blocks were executed verbatim from the .rst against the worktree with the assertion that gfql_explain reports used_index=True / index_selected for the Cypher hop. The docs example harness (shared namespace per file) passes on the page. The old claim that the native chain "is used automatically" is not carried forward: gfql_explain records no decision for that chain under either 'use' or 'force', so the page now points readers at used_index to check rather than asserting it. The multi-seed Cypher form (WHERE a.id IN [...]) scans instead of using the index; that gap is filed separately and the page leads with the single-seed form that is proven to engage. Opener reframed: GFQL is fast out of the box; an index is the usual next step when you or your coding agent want more. The unpublished-latency sentence and the three benchmark reproducer paths are removed, and Cost and fallback is three short bullets. No speed figure or competitor comparison is introduced; this page has no published latency for the index path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
pyg-bench #282 retired the IS3 single-query diagnostic probe, the one run that had aged past the 60-day policy. Nothing on any page referenced its two cells (both board_quotable=false), so no published number changes: 8 runs, 260 cells, same values. docs/test_bench_numbers.py goes from 1 failed to 38 passed. Provenance checked rather than assumed: the artifact's generated_by commit is an ancestor of pyg-bench main, and re-running scripts/export_docs_numbers.py on main reproduces the vendored file exactly apart from the timestamp stamps. The contract file is unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Two pages covered one topic from two entry points: index_adjacency.rst came first as the seeded-traversal guide, indexing.rst arrived three weeks later as the lifecycle guide when node_id/node_prop and gfql_index_all() landed, and each grew its own quick start and cost section while linking the other as "see also". The adjacency material is now one labelled subsection of the indexing guide (when to use it, build with Cypher, the index_policy table, column-stat facts); the duplicated quick start and cost text is gone. index_adjacency.rst stays as an :orphan: stub pointing at the new section, so index_adjacency.html keeps resolving for existing links (no redirect extension is installed and the build runs -W -n). All nine inbound :doc: references and the perf toctree entry are repointed; nothing else in docs/source references the old page. Two claims from the old text are not carried forward as written. "Measured figures are published on index_adjacency" pointed at a page with none. "Typed hops from known nodes (native chain or Cypher)" use the index: the Cypher single-seed forms do, verified via gfql_explain (used_index=True) with the edge type column bound; the IN-list form scans; the native chain records no decision, so the guide now says to check it with gfql_explain. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
The intro snippet and the Quick start now lead with Cypher strings: CREATE GFQL INDEX FOR edge_out_adj and node_id, SHOW GFQL INDEXES, a 1-hop and a 2-hop from a known node, decline safety via index_policy="off", and an on-page assert that gfql_explain reports used_index. The native chain and the direct hop() call follow as a secondary block. The seed-list form is named in prose as taking the scan path; no demo implies otherwise. The demos build the node-id index alongside the adjacency index. With the adjacency index alone the hop is served at the step level, but the top-level used_index that the page tells readers to check stays False because the destination-return seam wants node_id; building both makes the page's advice and the demo agree. The merged adjacency block gets the same treatment and its comment now states what explain reports rather than "served". Every block was executed verbatim from the .rst: each printed value matches its comment, used_index is True after exactly the two DDL lines, and the IN-list form reports used_index False. The chain result variable is renamed so it no longer shadows the Cypher result, and the staleness comment names only the indexes the Quick start now builds. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
… start
Alongside the row-returning MATCH ... RETURN, the quick start now shows the
same 1-hop and 2-hop lookups as GRAPH { MATCH ... } graph pipelines, whose
result keeps nodes and edges and can be plotted or queried further. Both forms
take the index path for a lookup from one known node; the prose says so and
covers the seed-list case for both: the row form scans, the GRAPH form does
not accept it yet.
Executed verbatim from the .rst with fresh-value asserts: the GRAPH results
are [0, 1, 2] with 2 edges and [0, 1, 2, 3, 4] with 5 edges, both
index_selected after exactly the two DDL lines, identical with indexes off,
and the GRAPH seed-list form raises GFQLValidationError as stated.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
A reader looking for a single-shot example will not find one, because none
exists: index DDL is accepted only as an entire standalone string, CREATE GFQL
INDEX is not a Cypher clause so it cannot appear inside GRAPH { }, and the
CreateIndex wire op is rejected as a chain stage or let binding. Rather than
leave the gap implicit, the quick start now states it in one sentence and
points at the tightest forms that do exist: g.gfql_index_all().gfql(query) and
the .gfql(ddl).gfql(query) chain already shown. The one-liner was executed
before being written down. The capability gap is filed separately.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
GFQL documentation now gives readers a short start path, separate performance and language-reference hubs, and an engine guide tied to measured results. Existing page URLs are preserved. The performance hub leads with the speedup case study and separates examples and architecture from benchmark tables and methodology.
The benchmark pages use verified publication cells for values, ratios, query tallies, and generated charts. This refresh includes GraphBench Q1–Q9 at 20k/100k from PyGraphistry
d20c6ae1af73e64b05aa4b24459dfd3ba247edf9and SNB SF0.1/SF1 fromf7a7253bc95d9cc0663bd130a53a03f2f355cbcd. The source revisions remain distinct. The final artifact is byte-identical to merged pyg-bench #267; #268 records and enforces matching CPU placement. Historical database measurements keep their original dates and methods.The engine guide explains input-based
autoselection, frame conversion, unsupported-query behavior, and when Polars-GPU still executes CPU operations. Indexed message-content and creator lookups lead the measured four-engine SNB comparisons at both scales. The full tables retain every eligible result. GraphBench 20k q8 is explicitly identified as CPU-only under Polars-GPU. Instrumented GPU diagnostics are excluded from published timings.The editorial pass addresses the recorded review items: shorter prose, defined lookup terms, consolidated engine decisions and memory guidance, simpler comparison tables, and one provenance section per page. The Neo4j/GDS and GraphFrames case studies retain their measured scope and limits. Benchmark contracts enforce source freshness, correct references, derived tallies, and chart synchronization.
Validation on the final data and prose:
$parameters in the case study; none point to the four final edited pages.Review previews:
Performance changes and private benchmark/TCK companions are merged. This public documentation PR is the remaining owner review and merge step. Final head
048b454673dcf673ba23784d938f9ab73845f02bpassed GitHub CI with 57 successful jobs and 11 conditional skips, plus CodeQL. Read the Docs build 34554961 passed on that exact head. Eleven hosted pages match the locally verified tables and headings; all six SVG assets load, and the final source/run identifiers were checked through the preview source link. The hosted performance page was also inspected in Chrome. The skipped route/TCK/GPU jobs do not add coverage; the separately recorded merged-master and private-companion validation remains the evidence for those areas.