SAH-KSE is an experimental semantic compression system that converts large datasets into compact knowledge seeds that preserve meaning while drastically reducing storage size.
Instead of compressing raw bytes like ZIP or Brotli, SAH-KSE compresses knowledge structure:
DATA → SEMANTIC ATOMS → CONCEPT GRAPH → ADAPTIVE HASH → KNOWLEDGE SEED
The resulting seed can be used for:
- 🧠 Low-RAM AI memory systems
- 🔎 Semantic search & retrieval
- 📦 Compact offline knowledge storage
- 🤖 LLM / VLM reasoning inputs
- 💻 CPU-only AI pipelines
Traditional compression:
File → smaller file
SAH-KSE:
Information → meaning → structure → tiny semantic representation
This enables:
- 10x–1000x reduction for structured knowledge
- Meaning-preserving storage instead of text storage
- Fast semantic lookup without loading full data
- Unified representation for text and images
Removes repeated or near-duplicate data using rolling hashes and similarity detection.
Breaks content into minimal meaning units:
Example:
Einstein developed relativity in 1905
Atoms:
(Einstein, developed, relativity)
(relativity, year, 1905)
Normalizes equivalent concepts into stable IDs:
Einstein → 1832
relativity → 771
developed → 44
Stores meaning as compact integer triples:
1832 44 771
771 7 1905
This representation is extremely space-efficient.
Generates a stable semantic fingerprint:
slot = atom_id % vector_size
vector[slot] += atom_id * prime_weight
hash_byte = vector % 256
Produces a 16–32 byte semantic hash that:
- is order-independent
- preserves semantic similarity
- runs using only integer math
- works efficiently on CPU
Final seed contains:
- concept dictionary
- compact knowledge graph
- semantic hashes
- metadata
Typical size:
KB–MB instead of GB
5MB article about photosynthesis.
Contains:
Concepts:
sunlight, chloroplast, CO2, water, glucose
Relations:
photosynthesis uses sunlight
photosynthesis produces glucose
chloroplast hosts photosynthesis
An LLM can regenerate explanations or answer questions using only this seed.
sah-kse/
│
├── sahkse/
│ ├── pipeline.py
│ ├── dedup.py
│ ├── atoms.py
│ ├── canonical.py
│ ├── graph.py
│ ├── sahash.py
│ ├── seed.py
│ └── reconstruct.py
│
├── examples/
├── tests/
├── README.md
├── requirements.txt
└── LICENSE
git clone https://github.com/<yourname>/sah-kse.git
cd sah-kse
pip install -r requirements.txtfrom sahkse.pipeline import compress_to_seed
seed = compress_to_seed("article.txt")
print(seed)from sahkse.seed import load_seed
seed = load_seed("article.seed")
print(seed.concepts)def semantic_hash(atom_ids, size=16):
primes = [3,5,7,11,13,17,19,23,29,31,37,41,43,47,53,59]
vec = [0]*size
for aid in atom_ids:
slot = aid % size
vec[slot] += aid * primes[slot]
return bytes([v % 256 for v in vec])- Text compression pipeline
- Semantic hashing module
- Seed file format
- Rust hashing core
- Streaming compression
- Faster canonicalization
- Image atom extraction
- Unified semantic space
- Tiny VLM support
- Seed-based retrieval
- LLM reconstruction
- Offline assistant demo
Contributions are welcome in:
- semantic hashing research
- compression experiments
- dataset benchmarking
- performance optimization
- Rust runtime modules
Open an issue or discussion to propose ideas.
MIT License (recommended for research & open collaboration)
SAH-KSE aims to become:
A foundational storage layer for AI-native systems where knowledge is stored as compact semantic seeds instead of raw files.
⭐ If this idea excites you, star the repo and help build the future of low-RAM AI memory systems.