This repository is the official implementation for the paper “The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA” by Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, and Stefano Rini. The preprint is available on arXiv:2609.27669.
The paper evaluates frozen, locally deployable small language models (SLMs) as question-conditioned graph-navigation policies, measuring terminal-answer accuracy together with executed-path fidelity. It also includes a Think-on-Graph (ToG)-style comparison to study the effect of explicit search and answer-generation scaffolding.
In the primary experimental setting, an SLM navigates the knowledge graph iteratively by selecting among the executable graph actions available at each step and deciding when to stop.
The repository supports three KGQA workflows:
- Iterative navigation QA: the LLM selects legal graph actions step by step. This is the paper's primary experimental setting.
- Think-on-Graph (ToG): a local adaptation of the original ToG search procedure used for the paper's secondary search comparison.
- Subgraph QA: the LLM receives a sampled subgraph and predicts the final answer directly.
The main entry points are kgqa_navigation.py, kgqa_tog.py, and kgqa_subgraph.py.
Python 3.12 or newer is required.
pip install -r requirements.txtThe experiments support either Ollama directly or Open WebUI backed by Ollama.
Copy the provided backend configuration template before running the experiments:
cp configs/openwebui_template.json configs/openwebui_config.jsonThen configure the selected backend and make sure the model referenced by your model profile is available.
See docs/backend.md for Ollama installation/model downloads, Open WebUI model availability, connection settings, and API-key setup.
Models are configured through validated JSON profiles under configs/models/ rather than repository-specific model-name conventions.
A profile records the backend model ID, model family, instruct/quantization metadata, context-window limit, and declared capabilities. Select one with:
--model-config configs/models/qwen3.jsonIf another server exposes the same deployed model under a different ID, override only the API-facing identifier:
--model-config configs/models/qwen3.json \
--model-id my-server-model-aliasSee configs/models/README.md for the profile schema and validation rules.
The paper evaluates the navigation-ready KINSHIP and MQuAKE-ST resources, including the Single Answer and Multi Answer MQuAKE-ST settings.
Dataset releases, preparation details, and associated THESEUS resources are maintained in the THESEUS repository. The processed datasets are also available from Hugging Face:
The preprocessing scripts expect the downloaded releases under raw_data/ and copy the files required by the experiment runners into data/. The scripts do not download the datasets themselves.
Using the Hugging Face CLI, the expected raw layouts can be created with:
hf download HalcyonSolutions/Kinship \
--repo-type dataset \
--local-dir ./raw_data/kinship_hinton
hf download HalcyonSolutions/MQuAKE-ST \
--repo-type dataset \
--local-dir ./raw_data/mquake_st_datasetEquivalent dataset downloads from Google Cloud Storage can be placed in the same raw_data/ directories.
Then preprocess both datasets from the repository root:
bash scripts/preprocess_kinship.sh
bash scripts/preprocess_mquake_st.shThis creates the runner-ready datasets under:
data/kinship/
data/mquake_single/
data/mquake_multi/
Both raw_data/ and data/ are local working directories and are excluded from version control.
See docs/datasets.md for the expected processed file layout and optional entity/relation mappings.
Run the experiment scripts from the repository root.
The one-shot navigation results and the zero-shot versus one-shot comparison are reproduced with:
bash scripts/kgqa_navigation_runs.shThe ToG-style search comparison is reproduced with:
bash scripts/kgqa_tog_runs.shThese scripts contain the model, dataset, prompting, navigation/search, and inference settings used for the reported experiments.
See docs/reproducibility.md for the exact configurations encoded by the scripts.
python ./kgqa_navigation.py \
--dataset mquake_single \
--hops n \
--model-config configs/models/qwen3.json \
--navigation-approach tuple \
--memory-approach full \
--prompting-approach zero-shot \
--n-shots 0 \
--max-navigation-steps 4 \
--max-actions 200 \
--result-dir ./resultsNavigation supports tuple, factorized, and hybrid action selection together with zero-shot and n-shot prompting.
See docs/navigation.md for prompting, memory, action-selection, truncation, context-window, and debugging options.
python ./kgqa_tog.py \
--dataset mquake_single \
--model-config configs/models/qwen3.jsonThe local adapter preserves the upstream ToG prompt family while adapting graph access, entity/relation mappings, evaluation, and result logging to this repository.
See docs/tog.md for the upstream revision, search procedure, fallback behavior, parsing policy, answer evaluation, implementation differences, audit trail, and result versioning.
python ./kgqa_subgraph.py \
--dataset mquake_single \
--hops n \
--model-config configs/models/qwen3.json \
--sampling-method neighborhood \
--subgraph-size 50 \
--max-depth 3 \
--result-dir ./resultsSee docs/subgraph.md for sampling and retrieval options.
analysis/ result compilation, diagnostics, comparisons, and plotting
configs/ API/backend configuration templates
docs/ detailed workflow and reproducibility documentation
model/ shared and task-specific LLM clients
scripts/ experiment, sample, sanity-check, and reproducibility scripts
tests/ unit, integration, and behavior tests
tools/ standalone debugging and development utilities
utils/ graph, metrics, KGQA, parsing, and API utilities
The local data/ directory and generated results/ directory are not checked into the repository.
For a module-level overview, see docs/development.md.
Generated outputs are written under results/.
Navigation results include the run configuration, aggregate statistics, per-question episodes, executed paths, path-fidelity and final-entity metrics, and prompt/action truncation metadata. ToG results additionally retain search, answer-source, parsing, and formatting audit information.
Common analysis commands include:
python analysis/compile_navigation_results.py --dataset kinship
python analysis/plot_metric.py --dataset mquakeSee docs/results.md for result fields, ToG audit metadata, and the available analysis utilities.
Detailed documentation is available under docs/:
If you use this repository, please cite:
@article{hernandez2026path,
title = {The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA},
author = {Eduin E. Hernandez and Sergio A. Diaz and Luis F. Garcia and Nurassyl Askar and Stefano Rini},
year = {2026},
journal = {arXiv preprint arXiv:2609.27669}
}This project is licensed under the Apache License 2.0.
Third-party datasets, model weights, and other external resources remain subject to their respective licenses and terms of use.
