SAGE (Suppression-Aware Graph Engine) is a production-oriented NLP and knowledge-graph reasoning system designed to answer questions over structured, temporally evolving data while preserving suppression semantics, provenance, temporal relationships, and deterministic graph reasoning.
SAGE combines Natural Language Processing (NLP), knowledge graphs, graph query generation, schema contracts, and suppression-aware reasoning into a controlled multi-stage pipeline.
The central design principle is:
Do not treat missing data as ordinary missing data when the source explicitly indicates that a value was suppressed, unavailable, confidential, below a reporting threshold, or otherwise non-observable.
SAGE therefore distinguishes between:
- an actual numerical value,
- an explicit zero,
- a suppressed value,
- an unavailable value,
- a missing value,
- and an unknown/unresolved value.
This distinction is critical for trustworthy analytical question answering.
- Overview
- Why SAGE Exists
- Core Problem
- What Makes SAGE Different
- Architecture
- Pipeline
- Repository Structure
- Core Contracts
- NLP Layer
- Knowledge Graph Layer
- Suppression Awareness
- Temporal Reasoning
- Query Generation
- Execution and Validation
- Technology Stack
- Getting Started
- Environment Setup
- Running SAGE
- Development
- Testing
- Production Design Principles
- Data Contracts
- Observability and Debugging
- Security Considerations
- Performance Considerations
- Design Trade-offs
- Roadmap
- Project Status
- Contributing
- License
Traditional question-answering systems generally follow a pattern similar to:
User Question
↓
LLM
↓
Generated Answer
This is convenient, but it is insufficient for analytical questions where correctness depends on:
- exact graph relationships,
- temporal ordering,
- entity resolution,
- explicit data semantics,
- missing-vs-suppressed distinctions,
- reproducible query execution,
- and validation of the generated result.
SAGE instead uses a controlled pipeline:
┌──────────────────────┐
│ User Query │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Query / NLP Layer │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Entity & Intent │
│ Resolution │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Semantic / Graph │
│ Planning │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Cypher Generation │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Graph Execution │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Result Validation │
│ + Suppression Logic │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Evidence / Answer │
└──────────────────────┘
The pipeline is deliberately separated into stages so that every transformation can be inspected, tested, validated, and reproduced.
Structured datasets frequently contain information that cannot safely be interpreted using ordinary SQL-style null semantics.
For example:
value = 0
does not necessarily mean the same thing as:
value = NULL
and neither necessarily means:
value = suppressed
Consider an emissions dataset:
January 10.4
February 11.2
March S
April 13.1
where S means that the source suppressed the value.
A naïve analytical system might transform this into:
January 10.4
February 11.2
March NULL
April 13.1
and subsequently calculate statistics or changes incorrectly.
SAGE preserves the distinction between observed, zero, suppressed, missing, and unknown states throughout the reasoning pipeline.
SAGE addresses analytical question answering over a knowledge graph where the answer depends on relationships rather than isolated records.
Typical questions include:
- "What was the emission for entity X in year Y?"
- "Which month had the highest value?"
- "What changed from one month to the next?"
- "Compare two entities."
- "What was January's change compared with December of the previous year?"
- "Which values were suppressed?"
- "What is the minimum observed value?"
- "Which records cannot be safely compared because the source suppressed a value?"
These questions require more than language understanding.
They require:
Language Understanding
+
Entity Resolution
+
Schema Understanding
+
Graph Traversal
+
Temporal Reasoning
+
Suppression Semantics
+
Deterministic Execution
+
Result Validation
SAGE is designed around that complete chain.
Suppressed observations are not silently converted into ordinary nulls or zeros.
Relationships are explicitly represented in the graph and traversed using Cypher.
Time relationships are modeled explicitly instead of relying only on string manipulation or application-side sorting.
Pipeline stages communicate through explicit schemas rather than undocumented Python dictionaries.
The LLM/NLP layer is separated from graph execution and validation.
Generated queries and returned results are checked before being exposed as final answers.
The system is designed to preserve where an answer came from and how it was derived.
Individual pipeline stages can be tested independently.
SAGE follows a layered architecture.
┌──────────────────────────────────────────────────────────┐
│ API Layer │
│ FastAPI / HTTP │
└───────────────────────────┬──────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────┐
│ Pipeline Manager │
│ Orchestrates the SAGE stages │
└───────────────────────────┬──────────────────────────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
NLP / Parsing Resolution Graph Planning
│ │ │
└──────────────┼──────────────┘
▼
Query Generation
│
▼
Graph Execution
│
▼
Validation / Semantics
│
▼
Answer Generation
The repository keeps contracts separate from agent/stage implementation.
This prevents individual components from independently redefining the data exchanged between stages.
SAGE is organized as a six-stage processing pipeline.
A conceptual representation is:
Stage 1
Input / Understanding
↓
Stage 2
Entity / Semantic Resolution
↓
Stage 3
Graph Planning
↓
Stage 4
Query Generation
↓
Stage 5
Execution
↓
Stage 6
Validation / Answer
The exact implementation of each stage may evolve, but the stage boundaries and contracts are treated as architectural interfaces.
This makes it possible to replace an implementation without breaking the rest of the system.
The core repository is organized around contracts, pipeline stages, API boundaries, and infrastructure.
SAGE---Suppression-Aware-Graph-Engine/
│
├── Dockerfile
├── Makefile
├── README.md
│
├── contracts/
│ ├── fact_schema.py
│ ├── stage_io.py
│ ├── CONTRACT_CHANGELOG.md
│ └── __pycache__/
│
├── src/
│ └── sage/
│ │
│ ├── api/
│ │ └── ...
│ │
│ ├── pipeline/
│ │ ├── manager.py
│ │ └── ...
│ │
│ ├── ...
│ │
│ └── ...
│
├── tests/
│ └── ...
│
└── ...
contracts/
is intentionally separate from the implementation.
Contracts define what stages are allowed to exchange.
Implementation modules consume those contracts.
SAGE uses explicit contracts for stage-to-stage communication.
Important contract modules include:
contracts/
├── fact_schema.py
├── stage_io.py
└── CONTRACT_CHANGELOG.md
Defines the canonical representation of extracted or resolved facts.
A fact can conceptually contain information such as:
entity
predicate
value
time
unit
status
provenance
confidence
The exact schema is defined by the implementation and should not be duplicated independently by downstream stages.
Defines the input/output boundaries between pipeline stages.
The objective is to prevent implicit coupling such as:
stage_a_output["something"]["nested"]["value"]without any contract guaranteeing that structure.
Instead, stage communication should be explicit and validated.
All intentional contract changes should be documented.
Contract changes are architectural changes because modifying a shared schema can affect multiple pipeline stages.
NLP is used where the system needs to convert natural-language intent into structured semantics.
For example:
"What was Comoros' NMVOC emission in March 2006?"
must be transformed into something conceptually equivalent to:
Entity:
Comoros
Pollutant:
NMVOC
Time:
March 2006
Measure:
Emission
Operation:
Retrieve
NLP therefore acts as the semantic interface between natural language and the graph.
It is not responsible for blindly producing the final answer.
SAGE uses a graph representation because the underlying problem is relationship-heavy.
A simplified graph may contain:
(:ReportingEntity)
│
│ HAS_SERIES
▼
(:EmissionSeries)
│
│ FOR_POLLUTANT
▼
(:Pollutant)
(:EmissionSeries)
│
│ HAS_MONTHLY_EMISSION
▼
(:MonthlyEmission)
│
│ NEXT_MONTH
▼
(:MonthlyEmission)
This structure allows questions involving relationships to be represented naturally.
For example:
January
│
│ NEXT_MONTH
▼
February
│
│ NEXT_MONTH
▼
March
This is materially different from merely storing:
month = 1
month = 2
month = 3
because the graph explicitly encodes the relationship between observations.
Suppression is a first-class semantic concept in SAGE.
The system must distinguish states such as:
| State | Meaning |
|---|---|
| Observed value | A usable reported numerical observation |
| Zero | Explicitly reported zero |
| Suppressed | Source intentionally hides or suppresses the value |
| Missing | No observation exists |
| Unknown | State cannot be established |
| Unavailable | Source indicates the observation is unavailable |
These states must not be collapsed prematurely.
For example:
0
and:
suppressed
are semantically different.
Similarly:
missing
does not automatically mean:
0
Temporal reasoning is a core part of SAGE.
The system should preserve temporal relationships rather than treating dates as ordinary text fields.
For example:
December 2005
│
│ NEXT_MONTH
▼
January 2006
│
│ NEXT_MONTH
▼
February 2006
Therefore, when a question asks:
Show each month's change from the month before.
January 2006 must be compared against:
December 2005
rather than:
January 2006 → no previous value
This is especially important for monthly time-series queries crossing year boundaries.
SAGE uses graph-query generation to translate resolved intent into executable Cypher.
A representative query may look like:
MATCH (e:ReportingEntity {code: 'COM'})
-[:HAS_SERIES]->(s:EmissionSeries)
-[:FOR_POLLUTANT]->(:Pollutant {code: 'NMVOC'})
MATCH (s)-[:HAS_MONTHLY_EMISSION]->(m:MonthlyEmission)
RETURN mFor temporal comparison, graph relationships can be traversed directly:
MATCH (current:MonthlyEmission)
<-[:HAS_MONTHLY_EMISSION]-(s:EmissionSeries)
-[:HAS_MONTHLY_EMISSION]->(previous:MonthlyEmission)
MATCH (previous)-[:NEXT_MONTH]->(current)
RETURN current, previousThe important principle is:
The generated query should express the user's semantic intent using the graph's canonical relationships.
It should not reconstruct graph semantics unnecessarily in application code.
Generated Cypher is not treated as the final authority.
The execution layer should:
- Validate the query.
- Execute it against the graph.
- Inspect the returned records.
- Preserve suppression state.
- Validate temporal relationships.
- Check expected cardinality.
- Detect ambiguous or incomplete results.
- Produce structured evidence.
- Only then generate the final answer.
Conceptually:
Natural Language
↓
Structured Intent
↓
Graph Plan
↓
Cypher
↓
Query Validation
↓
Neo4j
↓
Raw Results
↓
Semantic Validation
↓
Evidence
↓
Answer
SAGE is designed around a modern, production-oriented Python graph/NLP stack.
| Layer | Technology |
|---|---|
| Language | Python |
| API | FastAPI |
| Graph Database | Neo4j |
| Graph Query Language | Cypher |
| NLP / LLM Layer | Pluggable |
| Containerization | Docker |
| Development Environment | Linux / WSL2 compatible |
| Testing | Python testing ecosystem |
| Contracts | Typed Python schemas |
| Build / Tasks | Make |
The architecture intentionally avoids coupling the core reasoning engine to one specific LLM provider.
Install:
- Python 3.12+
- Docker
- Docker Compose, if using the containerized graph environment
- Git
- Neo4j
For Windows development, WSL2 with Ubuntu is supported and recommended for a Linux-compatible development environment.
git clone <REPOSITORY_URL>
cd SAGE---Suppression-Aware-Graph-EngineCreate a virtual environment:
python3 -m venv .venvActivate it:
source .venv/bin/activateUpgrade packaging tools:
python -m pip install --upgrade pipInstall project dependencies:
pip install -r requirements.txtIf the project uses an editable package installation:
pip install -e .Runtime configuration should be supplied through environment variables rather than hard-coded credentials.
Typical configuration includes:
NEO4J_URI=
NEO4J_USERNAME=
NEO4J_PASSWORD=
LLM_PROVIDER=
LLM_API_KEY=
LLM_MODEL=
Create a local environment file where appropriate:
cp .env.example .envNever commit secrets to Git.
If the repository provides a Docker configuration:
docker compose up -dCheck running containers:
docker psThe graph database should then be available to the SAGE application according to the configured Neo4j connection settings.
A typical FastAPI development command is:
uvicorn sage.api.main:app --reloadThe API will normally be available at:
http://localhost:8000
Interactive API documentation:
http://localhost:8000/docs
The exact entrypoint should follow the repository's current API module.
Where supported, common development operations can be exposed through the Makefile.
For example:
make helpand project-specific commands such as:
make test
make lint
make format
make runThe Makefile is the preferred place for standardized developer commands rather than requiring developers to memorize long command sequences.
A typical development workflow is:
1. Create / switch to a feature branch
2. Modify implementation
3. Validate contracts
4. Run unit tests
5. Run integration tests
6. Run static checks
7. Run the complete pipeline
8. Inspect generated Cypher
9. Validate graph results
10. Commit
Testing should exist at multiple levels.
Test individual components in isolation.
Examples:
Entity resolution
Date normalization
Suppression handling
Fact extraction
Query construction
Result normalization
Verify that pipeline stages respect the shared schemas.
Stage A output
↓
Contract validation
↓
Stage B input
This catches incompatible stage changes early.
Verify interaction between:
Pipeline
+
Neo4j
+
Cypher
An end-to-end test should exercise:
Natural-language question
↓
SAGE pipeline
↓
Graph query
↓
Graph execution
↓
Validation
↓
Final structured answer
SAGE is intended to be production-oriented, not merely a research prototype.
LLMs may be probabilistic, but critical graph operations should be deterministic.
For example:
Entity code resolution
Date arithmetic
Suppression interpretation
Graph traversal
Result validation
should not depend on free-form LLM reasoning.
Shared data structures should be explicitly defined.
Bad:
dict[str, object]for every stage.
Preferred:
Typed contract
↓
Validation
↓
Stage
↓
Typed output
The following responsibilities should remain separate:
NLP
Entity Resolution
Graph Planning
Query Generation
Graph Execution
Validation
Answer Formatting
This makes failures diagnosable.
If the system cannot establish that a result is valid, it should not silently invent certainty.
For example:
Suppressed
should not become:
0
just because the downstream calculation expects a number.
An analytical answer should ideally be traceable to:
User question
↓
Resolved entities
↓
Generated query
↓
Graph records
↓
Derived calculation
↓
Final answer
Contracts are architectural boundaries.
A contract change can affect:
Producer
↓
Contract
↓
Consumer
Therefore:
- Do not casually modify shared schemas.
- Document breaking changes.
- Update tests when contracts change.
- Keep compatibility considerations explicit.
- Record contract changes in
CONTRACT_CHANGELOG.md.
A production-grade pipeline should make failures inspectable.
Important debugging information includes:
Request ID
Pipeline stage
Resolved entities
Resolved temporal scope
Generated Cypher
Query parameters
Execution duration
Result cardinality
Suppression states
Validation failures
Final evidence
The goal is to make a failure answerable as:
Which stage produced the incorrect interpretation?
rather than simply:
The model gave the wrong answer.
SAGE should treat generated queries as untrusted output.
Recommended controls include:
- Parameterized Cypher where possible.
- Query validation before execution.
- Read-only graph credentials for analytical workloads.
- No secrets in source code.
- Environment-based configuration.
- Input validation at API boundaries.
- Request-level logging without leaking credentials.
- Restricted database permissions.
- Resource limits on expensive graph queries.
The LLM should never receive unrestricted authority to mutate the production graph.
SAGE's performance depends on both graph design and pipeline architecture.
Important considerations include:
Frequently queried properties should have appropriate Neo4j indexes.
Typical candidates include:
ReportingEntity.code
Pollutant.code
Prefer:
MATCH (e:ReportingEntity {code: $code})over unrestricted label scans.
Prefer:
MATCH (e:ReportingEntity {code: $code})with:
$code
rather than constructing Cypher through string concatenation.
For relationship-heavy operations, Cypher can often perform the required traversal more efficiently than retrieving large datasets into Python.
LLMs should be used for tasks that actually require language understanding.
Deterministic operations such as:
date arithmetic
sorting
aggregation
relationship traversal
schema validation
should remain deterministic.
SAGE intentionally chooses architectural control over minimal implementation complexity.
| Decision | Benefit | Cost |
|---|---|---|
| Multi-stage pipeline | Debuggability and modularity | More components |
| Explicit contracts | Safe interfaces | Schema maintenance |
| Neo4j | Natural relationship modeling | Operational complexity |
| Cypher | Powerful graph traversal | Requires graph-query expertise |
| Suppression-aware semantics | Correct analytical interpretation | More complex data model |
| LLM-assisted NLP | Natural-language interface | Probabilistic behavior |
| Deterministic validation | Reliability | Additional implementation |
| Provenance | Auditable answers | More metadata |
The objective is not to minimize lines of code.
The objective is to minimize silent semantic errors.
Consider:
"For Comoros and NMVOC in 2006, show each month's change from the month before."
SAGE should conceptually resolve:
Entity:
Comoros
Entity code:
COM
Pollutant:
NMVOC
Year:
2006
Operation:
Month-over-month change
Temporal rule:
January 2006 → December 2005
February 2006 → January 2006
...
December 2006 → November 2006
The important point is that January does not have a null predecessor merely because the requested year begins in January.
The graph's temporal relationship determines the correct predecessor.
A conceptual structure:
Comoros
│
│ HAS_SERIES
▼
EmissionSeries
│
│ FOR_POLLUTANT
▼
NMVOC
│
│ HAS_MONTHLY_EMISSION
▼
Dec 2005 ──NEXT_MONTH──> Jan 2006
│
└──NEXT_MONTH──> Feb 2006
This enables the system to calculate:
change(Jan 2006)
=
value(Jan 2006) - value(Dec 2005)
subject to suppression and observability rules.
SAGE should distinguish between:
No result
and:
Result exists but is suppressed
and:
Result exists but cannot be safely computed
and:
Query was ambiguous
and:
Query execution failed
These are different system states and should produce different responses.
SAGE follows several principles:
Do not calculate before establishing what the data means.
Do not reconstruct relationships unnecessarily outside the graph.
A suppressed observation is not equivalent to an absent observation.
Natural-language interpretation can be probabilistic.
Critical data operations should not be.
Pipeline stages should communicate through explicit interfaces.
The system should be able to establish:
What did the user ask?
What did we resolve?
What query did we execute?
What data did we retrieve?
How did we derive the result?
Potential future improvements include:
- Expanded suppression taxonomy
- Stronger provenance graph
- Automatic query-plan validation
- Query cost estimation
- More comprehensive temporal reasoning
- Improved entity resolution
- Confidence calibration
- Structured evidence generation
- Expanded benchmark suite
- Automated regression testing
- Production telemetry
- Distributed execution for large workloads
- Additional graph datasets
- Model-provider abstraction
- Retrieval and reasoning evaluation framework
SAGE is an actively developed engineering project.
The architecture prioritizes:
Correctness
>
Semantic integrity
>
Auditability
>
Maintainability
>
Performance
>
Convenience
Performance optimization should not compromise semantic correctness, particularly around suppression and temporal reasoning.
Contributions should preserve the architectural boundaries of the system.
Before submitting changes:
- Understand the relevant contract.
- Identify affected pipeline stages.
- Add or update tests.
- Update contract documentation for schema changes.
- Validate graph queries.
- Check suppression semantics.
- Check temporal edge cases.
- Verify that existing behavior remains correct.
For significant architectural changes, document the rationale before implementation.
Distributed under the Apache License 2.0. See LICENSE for more information.
SAGE is a suppression-aware, graph-native NLP reasoning engine.
Its purpose is not simply to convert:
Natural Language → Answer
but to build a controlled chain:
Natural Language
↓
Semantic Understanding
↓
Entity Resolution
↓
Graph Planning
↓
Cypher Generation
↓
Graph Execution
↓
Suppression-Aware Validation
↓
Temporal Reasoning
↓
Evidence
↓
Answer
The fundamental design goal is trustworthy analytical reasoning over structured temporal graph data.
SAGE — preserving the semantics that ordinary question-answering systems tend to lose.