Skip to content

Repository files navigation


⬡ ELIXARA

el·ixir + aura — the distilled essence, made luminous

Industrial Knowledge Intelligence · Built for the Engineers Who Keep the World Running


Every plant has a memory. Elixara is how it finally speaks.


License: MIT Platform: Windows CPU Only Zero Cloud Cost



The Name

An elixir, historically, was never just a liquid.

It was the product of a slow, deliberate process — raw matter, broken down, refined, reduced, until only the potent essence remained. Alchemists spent lifetimes on this. Not because they lacked intelligence, but because they believed that within ordinary things lived something extraordinary, waiting to be extracted.

Elixara takes that idea and gives it a hard drive.

Somewhere inside your plant's twenty thousand PDFs — the pump manuals no one rereads, the root-cause analysis filed and forgotten, the regulation clause buried on page 47 of a document that predates the current team — lives the actual intelligence of your operation. Not scattered. Not lost. Unextracted.

The -ara suffix carries the aura of that essence once it's freed: legible, searchable, citable, alive. Elixara doesn't manufacture knowledge. It distills what already exists, and gives it a voice.


The Problem, Stated Plainly

Industrial plants don't lack information. They drown in it.

A NASSCOM–EY study of Indian heavy manufacturing found an average of 7 to 12 disconnected systems per site, each speaking its own dialect, none listening to the others. McKinsey quantified the cost: 35% of a knowledge worker's time spent not doing the job, but finding what they need to do it.

Run the arithmetic on a plant with 1,000 engineers. 35% is roughly 350 engineer-years, lost annually — not to poor decisions, but to decisions made too slowly, or without the one document that would have made them obvious.

The clock is louder than the fragmentation. An estimated 25% of India's experienced industrial engineers retire within this decade. They carry tribal knowledge that was never written down twice — the "we tried that approach in 2019 and here's exactly why it failed" kind of knowledge. Unlike a server, a retiring engineer cannot be backed up.

Elixara exists because that knowledge cliff is not inevitable. It is a data architecture problem — and data architecture problems, given sufficient care, can be solved.


The Thesis

Treat every document in a plant as a neuron in a single industrial brain.

MinerU extracts the raw signal from the page. The chunker finds where one idea ends and the next begins, and injects structural context before either is ever embedded. nomic-embed-text draws the synaptic connections across 768 dimensions. ChromaDB holds the memory without needing a server. BM25 and dense vector retrieval, fused via Reciprocal Rank Fusion, vote on what's relevant. A cross-encoder reranker double-checks that vote with full joint attention over the query and each candidate passage. phi4-mini — running entirely on your CPU, no cloud, no API key, no monthly bill — synthesizes the relevant fragments into a cited, streamed, honest answer.

No single component here is exotic. The intelligence isn't in any one piece — it's in how deliberately they're wired together, and in the discipline to keep every claim traceable back to a real document, a real page, a real sentence.


What It Actually Does

Capability What it replaces What it feels like
Expert Copilot Searching 12 systems by hand Ask in plain English, receive a streamed, cited answer in seconds
Knowledge Graph Nobody's mental map of "what connects to what" A living, explorable web of equipment, regulations, people, and documents
Compliance Radar Manual audit prep, spreadsheet by spreadsheet Automated gap detection against OISD, PESO, and Factories Act clauses
Regulation Editor Static PDFs for standard operating procedures A MongoDB-backed editor to add, edit, and delete plant-specific rules in real-time
Failure DNA Tribal memory of "why does this keep breaking" A visual fingerprint of failure patterns per equipment class
Knowledge Coverage Heatmap Guessing what documentation is missing A live heatmap of document coverage across your entire equipment corpus
Preventive Maintenance Hunting through scattered manuals for schedules Automatic extraction of PM intervals into a unified, filterable dashboard
Ingestion Pipeline Someone manually tagging and filing PDFs Drop a file, watch it become structured, searchable knowledge in under two minutes

Every answer Elixara gives is grounded in a real excerpt, cited inline, with a confidence score attached. You can upvote or downvote answers to continuously calibrate that confidence over time. If the knowledge base genuinely doesn't contain the answer, Elixara says so — explicitly, every time. It does not guess at torque values. It does not paraphrase a safety regulation. It would rather admit ignorance than invent a hazard.


System Architecture

graph TD
    Client[Web Client<br/>React 18 + Vite] <-->|HTTP / SSE| Gateway[API Gateway<br/>Node + Express :4000]

    Gateway <-->|/docs · /upload · /jobs| Ingest[Ingest Service<br/>FastAPI :5001]
    Gateway <-->|/query · /stream| RAG[RAG Service<br/>FastAPI :5002]
    Gateway <-->|/nodes · /edges · /path| Graph[Graph Service<br/>FastAPI :5003]
    Gateway <-->|/scan · /report| Compliance[Compliance Service<br/>FastAPI :5004]

    Ingest -->|Parse + OCR| MinerU[MinerU CLI<br/>CPU pipeline backend]
    Ingest -->|Write metadata + entities| Mongo[(MongoDB)]
    Ingest -->|Store 768-dim vectors| Chroma[(ChromaDB)]

    RAG -->|Cache-aside · TTL 1h| Redis[(Memurai / Redis)]
    RAG -->|Read chunks| Mongo
    RAG -->|Cosine search| Chroma
    RAG -->|Generate + Embed| Ollama[Ollama<br/>phi4-mini · nomic-embed-text]

    Graph -->|BFS + CRUD| Mongo
    Compliance -->|Cross-check each clause| RAG
Loading

The data flow, in one breath:

A PDF lands on disk → MinerU produces clean Markdown and a structural map → the chunker respects document headings, injects hierarchical context into each segment before it's embedded → phi4-mini extracts equipment tags, regulation IDs, personnel names, and dates in a single JSON-mode pass → nomic-embed-text encodes each chunk as a 768-dimensional vector using the search_document: prefix that preserves retrieval accuracy most pipelines quietly sacrifice → the graph builder upserts every named entity as a node and every co-occurrence as a weighted edge → the document is now queryable, connected, and cited. Permanently.


Technology Stack — and Why Each Piece

Every choice below was made against one hard constraint: runs on a CPU laptop, costs nothing to operate, never phones home.

Layer Technology Why this one, specifically
Document parsing MinerU (pipeline backend) 86.2 on OmniDocBench v1.5 — layout-aware, table-aware, formula-aware, CPU-only, fully local
Local LLM phi4-mini via Ollama 3.8B params, native JSON mode, 128K context — small enough to run alongside everything else, sharp enough to follow strict citation rules without hallucinating
Embeddings nomic-embed-text 768-dim, 8,192-token context window — and critically, the search_document: / search_query: prefix pairing that recovers 5–8% retrieval accuracy most implementations silently discard
Vector store ChromaDB (embedded) Zero ops overhead — no server process, no port, persists straight to disk
Sparse retrieval BM25Okapi (rank-bm25) Catches exact-term matches ("torque specification," "OISD-118 clause 7.4") that pure semantic search misses
Score fusion Reciprocal Rank Fusion, k=60 Merges dense + sparse rankings without requiring comparable raw score distributions
Reranking ms-marco-MiniLM-L-6-v2 22MB cross-encoder — joint query-passage encoding for a precision pass that first-stage retrieval cannot afford
Primary database MongoDB Schema flexibility for heterogeneous document metadata; native $push / $addToSet / $inc accumulation patterns for the knowledge graph
Cache Memurai (Redis-compatible, Windows-native) Cache-aside on query results with 1h TTL — repeated demo queries return in single-digit milliseconds
Gateway Node.js + Express Non-blocking event loop is the right fit for SSE pipe-through and multipart file uploads
Frontend React 18 + Vite + Tailwind + Zustand + D3 Streaming UI, force-directed knowledge graph, and a design system built to feel less like a dashboard and more like a place you'd actually want to spend eight hours

Project Structure

elixara/
├── .env.example                      ← copy to .env before first run
├── package.json                      ← root orchestration (Windows-adapted)
│
├── frontend/                         ← React 18 SPA
│   └── src/
│       ├── pages/                      12 routes: dashboard · upload · query · graph ·
│       │                               equipment · failures · compliance · maintenance ·
│       │                               analytics · settings · library · login
│       ├── components/                 layout · ui primitives · docs · query · graph · charts
│       ├── store/                      Zustand slices: app · docs · query · graph
│       ├── hooks/                      ingestion polling · SSE streaming · D3 lifecycle
│       └── api/                        thin axios wrappers per domain
│
├── gateway/                          ← Node.js API Gateway :4000
│   ├── routes/                         docs · query · graph · compliance · health · auth
│   └── middleware/                     JWT auth · multer upload · morgan logging
│
└── services/                         ← Python microservices
    ├── shared/                         config · mongo · chroma · ollama_client · models
    ├── ingest_service/    :5001        MinerU → chunk → NER → embed → graph build
    ├── rag_service/       :5002        dense + BM25 + RRF → rerank → stream
    ├── graph_service/     :5003        node/edge CRUD · BFS pathfinder
    └── compliance_service/ :5004       regulation DB · RAG-backed gap scan

Prerequisites

Install these once, before touching the project folder. MongoDB, Memurai, and Ollama each register as background Windows services or processes — you will not restart them manually every session.

Tool Purpose Install from Verify with
Node.js 20+ Frontend + Gateway nodejs.org node --version
Python 3.11 AI microservices python.org python --version
Git Version control git-scm.com git --version
MongoDB Community Primary database mongodb.com/try/download/community — .msi, keep Install as a Service checked mongosh --version
Memurai Developer Edition Redis-compatible cache memurai.com/get-memurai memurai-cli ping → PONG
Ollama Local LLM runtime ollama.com ollama list

Pull both required models once (~2.7GB total):

ollama pull phi4-mini
ollama pull nomic-embed-text

On where you build this. Avoid hosting the project inside a OneDrive-synced folder. OneDrive attempts to sync every file the moment venv or node_modules creates it — producing intermittent file-lock errors during pip install. If you hit unexplained permission errors, move the project to C:\dev\elixara and the problem disappears.


Complete Setup — Windows / PowerShell, End to End

Every command below assumes you're in the elixara\ project root, in a PowerShell terminal. VS Code's integrated terminal is fine.

Step 0 — Clone and enter the project

git clone <your-repo-url> elixara
cd elixara

Step 1 — Environment variables

Copy-Item .env.example .env

The defaults already point at localhost for MongoDB, Memurai, and Ollama. No edits required for local development.

Step 2 — Install MinerU in its own isolated environment

MinerU carries heavy dependencies (PaddleOCR, ONNX Runtime), so it lives in a dedicated environment outside the project's own venvs:

python -m venv $HOME\.mineru_env
C:\Users\<you>\.mineru_env\Scripts\Activate.ps1

Use the full explicit path rather than $HOME when activating — the $HOME-prefixed form can pick up invisible characters on copy-paste in PowerShell that silently break the command. If PowerShell refuses with "execution of scripts is disabled", run this once and retry: Set-ExecutionPolicy RemoteSigned -Scope CurrentUser

pip install "mineru[core]"

MinerU 3.x downloads its parsing models automatically on first real use. Trigger that download now, deliberately, with a real PDF:

mineru -p "C:\path\to\any\test.pdf" -o C:\temp\mineru_test -b pipeline

The first run pulls roughly 1GB of models (layout detection, OCR, table recognition) and can take 5–15 minutes. Every subsequent run is fast. Confirm it worked:

Get-ChildItem -Recurse C:\temp\mineru_test -Filter *.md

Output structure note. MinerU 3.4 nests output one level deeper than older documentation describes — <output_dir>\<filename>\auto\<filename>.md, not directly under <filename>\. The codebase already searches recursively with rglob("*.md"), so no code change is needed.

Exit the environment and add MinerU permanently to your PATH:

deactivate
[Environment]::SetEnvironmentVariable("Path", $env:Path + ";$HOME\.mineru_env\Scripts", "User")

Fully close and reopen VS Code — not just a new terminal tab. PATH changes only take effect after a full application restart. Then verify:

mineru --version

Step 3 — Install Node dependencies

npm install

cd frontend && npm install && cd ..
cd gateway  && npm install && cd ..

Step 4 — Loosen two dependency pins before installing Python services

MinerU 3.4 pulls in gradio as an optional dependency, which requires newer fastapi and python-multipart than the original pins. Open services\shared\requirements.txt and confirm:

fastapi>=0.115.2
python-multipart>=0.0.18

If they still show exact pins, loosen them — otherwise pip install will fail with a ResolutionImpossible error.

Step 5 — Create a virtual environment for each Python microservice

ingest_service:

cd services\ingest_service
python -m venv .venv
.\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..

rag_service — install CPU-only PyTorch before the requirements file, or sentence-transformers will pull a CUDA build:

cd services\rag_service
python -m venv .venv
.\.venv\Scripts\pip.exe install torch==2.3.1+cpu --index-url https://download.pytorch.org/whl/cpu
.\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..

graph_service and compliance_service:

cd services\graph_service
python -m venv .venv && .\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..

cd services\compliance_service
python -m venv .venv && .\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..

Step 6 — Confirm MongoDB, Memurai, and Ollama are running

mongosh --eval "db.runCommand({ ping: 1 })"
memurai-cli ping
ollama list

If Ollama shows no models, pull phi4-mini and nomic-embed-text as listed in Prerequisites. If Memurai doesn't respond, open services.msc and confirm the Memurai service is set to Running.


Running Elixara

From the project root, one command starts everything — frontend, gateway, and all four Python microservices, concurrently:

npm run dev:all

Give it a few seconds. You're looking for six clean startup lines with no traceback beneath any of them:

[dev:frontend]   VITE ready
[dev:gateway]    ⬡  Elixara Gateway running on :4000
[dev:ingest]     Uvicorn running on http://127.0.0.1:5001
[dev:rag]        Uvicorn running on http://127.0.0.1:5002
[dev:graph]      Uvicorn running on http://127.0.0.1:5003
[dev:compliance] Uvicorn running on http://127.0.0.1:5004

Then open:

Demo login: demo / elixara2024 — or hit ⚡ Judge Demo Access for an instant token.

Seed demo documents (optional)

python scripts\seed_demo.py

Uploads sample PDFs through the full ingestion pipeline. If demo_docs\ doesn't exist yet, add your own PDFs first — or drag files directly onto the Upload page in the running app.


Verifying Everything Is Alive

python scripts\check_health.py

Or open the Settings page inside the app (http://localhost:5173/settings) — it live-polls every service and Ollama itself, showing a green or red dot with latency for each.

A fully healthy stack:

{
  "ingest":     { "status": "ok" },
  "rag":        { "status": "ok" },
  "graph":      { "status": "ok" },
  "compliance": { "status": "ok" },
  "ollama":     { "status": "ok" }
}

Services & Ports Reference

Service Framework Port Interactive docs
Frontend React 18 / Vite 5173 http://localhost:5173
API Gateway Express 4000 http://localhost:4000/api/*
Ingest Service FastAPI 5001 http://localhost:5001/docs
RAG Service FastAPI 5002 http://localhost:5002/docs
Graph Service FastAPI 5003 http://localhost:5003/docs
Compliance Service FastAPI 5004 http://localhost:5004/docs
Ollama — 11434 http://localhost:11434
MongoDB — 27017 Windows Service
Memurai — 6379 Windows Service

Troubleshooting

Symptom Root cause Fix
mongosh not recognized MongoDB not installed, or terminal opened before install Reinstall via .msi with Install as a Service checked; open a fresh terminal
memurai-cli not recognized Same, for Memurai Reinstall; confirm the Memurai service is Running in services.msc
mineru not recognized after setting PATH VS Code caches environment variables per-window Fully close all VS Code windows and reopen — a new terminal tab alone is not enough
PowerShell parser error on $HOME\.mineru_env\... Invisible characters from copy-paste Retype manually or use the full explicit path instead of $HOME
ResolutionImpossible on fastapi or python-multipart MinerU's gradio dependency requires newer versions than the original pins In services/shared/requirements.txt: loosen to fastapi>=0.115.2 and python-multipart>=0.0.18
TypeError: unsupported operand type(s) for | in chroma.py chromadb.PersistentClient is a factory, not a class — the X | None type union fails at import time Add from __future__ import annotations as the first import in services/shared/chroma.py
NameError: name 'logging' is not defined in scan.py Missing standard-library import Add import logging at the top of compliance_service/routers/scan.py
"The system cannot find the path specified" for dev:graph or dev:compliance That service's .venv was never created Run Step 5 for the missing service
.venv/bin/uvicorn errors Copy-pasted Linux path syntax Windows uses .venv\Scripts\uvicorn, not .venv/bin/uvicorn
pip install fails on torch Wrong index URL or missing CPU flag Use exactly: pip install torch==2.3.1+cpu --index-url https://download.pytorch.org/whl/cpu
Port already in use A previous process didn't shut down cleanly Get-Process -Id (Get-NetTCPConnection -LocalPort 5001).OwningProcess | Stop-Process
SSE answers appear all at once instead of token-by-token Proxy buffering Confirm res.flushHeaders() runs and X-Accel-Buffering: no is set — already configured in this repo's gateway
ChromaDB "collection not found" Corrupted local vector store Delete data\chroma_db\, restart the RAG service, re-ingest
Symlink warning during MinerU model download Windows blocks symlinks without Developer Mode Harmless — Hugging Face falls back to full file copies; no functional impact
Intermittent file-permission errors during dependency install Project inside a OneDrive-synced folder Move to C:\dev\elixara

A Closing Note

There is something almost stubborn about the idea at the center of this project.

Not that the problem is hard — though it is. Not that the technology is novel — though it is assembled carefully. But that the answer was already there. Not behind a paywall. Not requiring a new sensor, a new survey, a new budget line. Buried, instead, under the ordinary weight of twelve disconnected systems and a PDF nobody has opened since the inspection that produced it.

I built Elixara because I kept thinking about those retiring engineers — the ones who carry in their heads the reason a particular pump keeps failing, the way a particular regulation intersects with a particular equipment class, the shortcut that saved three hours the last time this exact alarm fired. When they leave, that knowledge doesn't get archived. It evaporates.

Elixara doesn't claim to know more than your engineers do. It claims to remember what they already knew, and to say it back the moment someone needs it — with the receipt attached, every time, so that trust is never asked for without being earned.

The plant already has a brain. Elixara is how it finally speaks.


License

Distributed under the MIT License. See LICENSE for details.


⬡ Elixara · CPU-Only · Zero Cloud Cost · Zero Hallucination Tolerance

Built for the industrial engineers of tomorrow — and the knowledge they carry today.

About

CPU-only industrial knowledge intelligence: hybrid RAG + knowledge graph + compliance automation for plant documents. No cloud, no API costs.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages