Skip to content

Commit 28d181c

Browse files
authored
ci: add codespell spell check and extend python linting to dev scripts (#864)
* ci: add codespell spell check and extend python linting to dev scripts - Add codespell (config in .codespellrc) covering yml, python and scala via a new Spell Check workflow and a pre-commit hook; fix the typos it surfaced across sources and docs. graphx/ is excluded to stay aligned with upstream Apache Spark. - Extend black/flake8/isort to the dev/ and python/dev/ scripts in both the Python CI code-style step and pre-commit. * ci: pass explicit config to black/isort so root dev/ does not break discovery When ../dev is included, black/isort resolve their config from the common parent of all paths (the repo root, which has no pyproject.toml) and fall back to defaults. Pass --config/--settings-path explicitly. * ci: format dev connect scripts to satisfy extended linting run_connect.py and stop_connect.py were added in #863 before the linting scope was extended to dev/. Apply black/isort and add a noqa for an unavoidable long config string so python-ci passes.
1 parent f4b2654 commit 28d181c

37 files changed

Lines changed: 99 additions & 69 deletions

‎.codespellrc‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,12 @@
1+
[codespell]
2+
# Skip generated, vendored, binary, and lock files.
3+
# graphx/ is ported from Apache Spark GraphX; keep it aligned with upstream.
4+
# graphframes-connect-databricks/ holds local sbt build output (gitignored).
5+
skip = .git,*.lock,poetry.lock,*.iml,target,build,dist,*_pb2*.py,*_pb2_grpc.py,graphx,graphframes-connect-databricks,*.svg,*.png,*.gif,*.parquet,*.csv,*.json,*.pom,*.semanticdb
6+
# Identifiers and domain terms that are not misspellings:
7+
# mis/MIS - Maximal Independent Set
8+
# te/ser - local variable names in Scala tests/algorithms
9+
# fof/mye/protocall/alledges - local variable names (camelCase) in tests
10+
# afterall - ScalaTest `afterAll` lifecycle override
11+
# jurney - maintainer surname (Russell Jurney)
12+
ignore-words-list = mis,te,ser,fof,mye,protocall,alledges,afterall,jurney

‎.github/workflows/python-ci.yml‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -57,9 +57,9 @@ jobs:
5757
- name: Code style
5858
working-directory: ./python
5959
run: |
60-
poetry run python -m black --check graphframes
61-
poetry run python -m flake8 graphframes
62-
poetry run python -m isort --check graphframes
60+
poetry run python -m black --check --config pyproject.toml graphframes dev ../dev
61+
poetry run python -m flake8 graphframes dev ../dev
62+
poetry run python -m isort --check --settings-path pyproject.toml graphframes dev ../dev
6363
- name: Build jar
6464
working-directory: ./python
6565
run: |

‎.github/workflows/spellcheck.yml‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,17 @@
1+
name: Spell Check
2+
on: [push, pull_request]
3+
permissions:
4+
contents: read
5+
jobs:
6+
codespell:
7+
runs-on: ubuntu-latest
8+
steps:
9+
- uses: actions/checkout@v6
10+
- uses: actions/setup-python@v6
11+
with:
12+
python-version: "3.12"
13+
- name: Install codespell
14+
run: pip install codespell==2.4.2
15+
- name: Run codespell
16+
# Configuration (skips and ignore list) lives in .codespellrc.
17+
run: codespell .

‎.pre-commit-config.yaml‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -3,19 +3,19 @@ repos:
33
hooks:
44
- id: black
55
name: black
6-
entry: bash -c 'cd python && poetry run black graphframes/' --
6+
entry: bash -c 'cd python && poetry run black --config pyproject.toml graphframes/ dev/ ../dev/' --
77
language: system
88
types: [python]
99

1010
- id: flake8
1111
name: flake8
12-
entry: bash -c 'cd python && poetry run flake8 graphframes/' --
12+
entry: bash -c 'cd python && poetry run flake8 graphframes/ dev/ ../dev/' --
1313
language: system
1414
types: [python]
1515

1616
- id: isort
1717
name: isort
18-
entry: bash -c 'cd python && poetry run isort graphframes/' --
18+
entry: bash -c 'cd python && poetry run isort --settings-path pyproject.toml graphframes/ dev/ ../dev/' --
1919
language: system
2020
types: [python]
2121

@@ -32,3 +32,9 @@ repos:
3232
language: system
3333
types: [scala]
3434
pass_filenames: false
35+
36+
- repo: https://github.com/codespell-project/codespell
37+
rev: v2.4.2
38+
hooks:
39+
- id: codespell
40+
# Configuration (skips and ignore list) lives in .codespellrc.

‎CONTRIBUTING.md‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -117,7 +117,8 @@ Enhancement suggestions are tracked as [GitHub issues](https://github.com/graphf
117117
We are using the following tools for enforcing of the codestyle:
118118

119119
- for python we are relying on the `black` + `flake8` + `isort`;
120-
- for scala we are relying on the `scalafmt`.
120+
- for scala we are relying on the `scalafmt`;
121+
- for spelling (across yml, python, and scala) we are relying on `codespell` (configured in `.codespellrc`).
121122

122123
You can enforce the code style using the [pre-commit-hooks](https://pre-commit.com/):
123124

‎README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ There are some popular use cases when GraphFrames is almost irreplaceable, inclu
1717
- Compliance analytics with a scalable shortest paths algorithm and motif analysis;
1818
- Anti-fraud with scalable cycles detection in large networks and by using K-Core algorithm;
1919
- Identity resolution at the scale of billions with highly efficient connected components;
20-
- Plan marketing campaigns in social networks using Maximal Indpendent Set algorithm;
20+
- Plan marketing campaigns in social networks using Maximal Independent Set algorithm;
2121
- Rank search result with a distributed, Pregel-based PageRank;
2222
- Cluster huge graphs with Label Propagation and Power Iteration Clustering;
2323
- Compute node embeddings at billion scale using Random-Walks and Hash2Vec model;

‎core/src/main/scala/org/graphframes/GraphFrame.scala‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -623,7 +623,7 @@ class GraphFrame private (
623623
case VarLengthPattern(src, name, min, max, direction, dst) =>
624624
if (min.isEmpty || max.isEmpty) {
625625
throw new InvalidParseException(
626-
s"Unbounded length patten ${pattern} is not supported! " +
626+
s"Unbounded length pattern ${pattern} is not supported! " +
627627
"Please a pattern of defined length.")
628628
}
629629
findVarLengthPattern(src, name, min.toInt, max.toInt, direction, dst)
@@ -982,7 +982,7 @@ class GraphFrame private (
982982
* large-scale sparse graphs." Proceedings of Simpósio Brasileiro de Pesquisa Operacional
983983
* (SBPO’15) (2015): 1-11.
984984
*
985-
* Returns a DataFrame with unque cycles.
985+
* Returns a DataFrame with unique cycles.
986986
*
987987
* @return
988988
* an instance of DetectingCycles initialized with the current context

‎core/src/main/scala/org/graphframes/embeddings/RandomWalkEmbeddings.scala‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -230,7 +230,7 @@ class RandomWalkEmbeddings private[graphframes] (private val graph: GraphFrame)
230230
.run()
231231
} else {
232232
// we need to create a new DataFrame, so unpersisting peristedEmbeddings
233-
// does not accidently unpersist the result;
233+
// does not accidentally unpersist the result;
234234
// dummy operations are cheap, but will create a different plan.
235235
persistedEmbeddings
236236
.withColumnRenamed(RandomWalkEmbeddings.embeddingColName, "x")
@@ -269,7 +269,7 @@ object RandomWalkEmbeddings extends Serializable {
269269
* provide a smooth way to initialize the whole embeddings pipeline with a single method call
270270
* that is usable for Python API (py4j and Spark Connect).
271271
*
272-
* Instead of this API it is recommended to use new + settters of the class!
272+
* Instead of this API it is recommended to use new + setters of the class!
273273
*/
274274
def pythonAPI(
275275
graph: GraphFrame,

‎core/src/main/scala/org/graphframes/lib/AllPaths.scala‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -49,7 +49,7 @@ import org.graphframes.WithLocalCheckpoints
4949
* - `len`: number of edges in the path (Long)
5050
*
5151
* Note: in the case of undirected graph an algorithm run on the internal graph made by union
52-
* edges and reversed edges. It is assummed that graph does not have multi-edges. Results may be
52+
* edges and reversed edges. It is assumed that graph does not have multi-edges. Results may be
5353
* unstable and unpredictable for the graph with multi-edges.
5454
*/
5555
class AllPaths private[graphframes] (private val graph: GraphFrame)

‎core/src/main/scala/org/graphframes/lib/HyperANF.scala‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -171,7 +171,7 @@ class HyperANF private[graphframes] (graph: GraphFrame)
171171
.groupBy(col(GraphFrame.SRC).alias(GraphFrame.ID))
172172
.agg(hll_union_agg(s"hop_${hop - 1}").alias(s"hop_${hop}"))
173173

174-
// stanard GF persist-unpersist-checkpoint flow
174+
// standard GF persist-unpersist-checkpoint flow
175175
state = {
176176
val stateToPersist = state.join(nState, GraphFrame.ID)
177177
if (shouldCheckpoint && hop % checkpointInterval == 0) {

0 commit comments

Comments
 (0)