Industrial Knowledge Intelligence · Built for the Engineers Who Keep the World Running
Every plant has a memory. Elixara is how it finally speaks.
An elixir, historically, was never just a liquid.
It was the product of a slow, deliberate process — raw matter, broken down, refined, reduced, until only the potent essence remained. Alchemists spent lifetimes on this. Not because they lacked intelligence, but because they believed that within ordinary things lived something extraordinary, waiting to be extracted.
Elixara takes that idea and gives it a hard drive.
Somewhere inside your plant's twenty thousand PDFs — the pump manuals no one rereads, the root-cause analysis filed and forgotten, the regulation clause buried on page 47 of a document that predates the current team — lives the actual intelligence of your operation. Not scattered. Not lost. Unextracted.
The -ara suffix carries the aura of that essence once it's freed: legible, searchable, citable, alive. Elixara doesn't manufacture knowledge. It distills what already exists, and gives it a voice.
Industrial plants don't lack information. They drown in it.
A NASSCOM–EY study of Indian heavy manufacturing found an average of 7 to 12 disconnected systems per site, each speaking its own dialect, none listening to the others. McKinsey quantified the cost: 35% of a knowledge worker's time spent not doing the job, but finding what they need to do it.
Run the arithmetic on a plant with 1,000 engineers. 35% is roughly 350 engineer-years, lost annually — not to poor decisions, but to decisions made too slowly, or without the one document that would have made them obvious.
The clock is louder than the fragmentation. An estimated 25% of India's experienced industrial engineers retire within this decade. They carry tribal knowledge that was never written down twice — the "we tried that approach in 2019 and here's exactly why it failed" kind of knowledge. Unlike a server, a retiring engineer cannot be backed up.
Elixara exists because that knowledge cliff is not inevitable. It is a data architecture problem — and data architecture problems, given sufficient care, can be solved.
Treat every document in a plant as a neuron in a single industrial brain.
MinerU extracts the raw signal from the page. The chunker finds where one idea ends and the next begins, and injects structural context before either is ever embedded. nomic-embed-text draws the synaptic connections across 768 dimensions. ChromaDB holds the memory without needing a server. BM25 and dense vector retrieval, fused via Reciprocal Rank Fusion, vote on what's relevant. A cross-encoder reranker double-checks that vote with full joint attention over the query and each candidate passage. phi4-mini — running entirely on your CPU, no cloud, no API key, no monthly bill — synthesizes the relevant fragments into a cited, streamed, honest answer.
No single component here is exotic. The intelligence isn't in any one piece — it's in how deliberately they're wired together, and in the discipline to keep every claim traceable back to a real document, a real page, a real sentence.
| Capability | What it replaces | What it feels like |
|---|---|---|
| Expert Copilot | Searching 12 systems by hand | Ask in plain English, receive a streamed, cited answer in seconds |
| Knowledge Graph | Nobody's mental map of "what connects to what" | A living, explorable web of equipment, regulations, people, and documents |
| Compliance Radar | Manual audit prep, spreadsheet by spreadsheet | Automated gap detection against OISD, PESO, and Factories Act clauses |
| Regulation Editor | Static PDFs for standard operating procedures | A MongoDB-backed editor to add, edit, and delete plant-specific rules in real-time |
| Failure DNA | Tribal memory of "why does this keep breaking" | A visual fingerprint of failure patterns per equipment class |
| Knowledge Coverage Heatmap | Guessing what documentation is missing | A live heatmap of document coverage across your entire equipment corpus |
| Preventive Maintenance | Hunting through scattered manuals for schedules | Automatic extraction of PM intervals into a unified, filterable dashboard |
| Ingestion Pipeline | Someone manually tagging and filing PDFs | Drop a file, watch it become structured, searchable knowledge in under two minutes |
Every answer Elixara gives is grounded in a real excerpt, cited inline, with a confidence score attached. You can upvote or downvote answers to continuously calibrate that confidence over time. If the knowledge base genuinely doesn't contain the answer, Elixara says so — explicitly, every time. It does not guess at torque values. It does not paraphrase a safety regulation. It would rather admit ignorance than invent a hazard.
graph TD
Client[Web Client<br/>React 18 + Vite] <-->|HTTP / SSE| Gateway[API Gateway<br/>Node + Express :4000]
Gateway <-->|/docs · /upload · /jobs| Ingest[Ingest Service<br/>FastAPI :5001]
Gateway <-->|/query · /stream| RAG[RAG Service<br/>FastAPI :5002]
Gateway <-->|/nodes · /edges · /path| Graph[Graph Service<br/>FastAPI :5003]
Gateway <-->|/scan · /report| Compliance[Compliance Service<br/>FastAPI :5004]
Ingest -->|Parse + OCR| MinerU[MinerU CLI<br/>CPU pipeline backend]
Ingest -->|Write metadata + entities| Mongo[(MongoDB)]
Ingest -->|Store 768-dim vectors| Chroma[(ChromaDB)]
RAG -->|Cache-aside · TTL 1h| Redis[(Memurai / Redis)]
RAG -->|Read chunks| Mongo
RAG -->|Cosine search| Chroma
RAG -->|Generate + Embed| Ollama[Ollama<br/>phi4-mini · nomic-embed-text]
Graph -->|BFS + CRUD| Mongo
Compliance -->|Cross-check each clause| RAG
The data flow, in one breath:
A PDF lands on disk → MinerU produces clean Markdown and a structural map → the chunker respects document headings, injects hierarchical context into each segment before it's embedded → phi4-mini extracts equipment tags, regulation IDs, personnel names, and dates in a single JSON-mode pass → nomic-embed-text encodes each chunk as a 768-dimensional vector using the search_document: prefix that preserves retrieval accuracy most pipelines quietly sacrifice → the graph builder upserts every named entity as a node and every co-occurrence as a weighted edge → the document is now queryable, connected, and cited. Permanently.
Every choice below was made against one hard constraint: runs on a CPU laptop, costs nothing to operate, never phones home.
| Layer | Technology | Why this one, specifically |
|---|---|---|
| Document parsing | MinerU (pipeline backend) | 86.2 on OmniDocBench v1.5 — layout-aware, table-aware, formula-aware, CPU-only, fully local |
| Local LLM | phi4-mini via Ollama | 3.8B params, native JSON mode, 128K context — small enough to run alongside everything else, sharp enough to follow strict citation rules without hallucinating |
| Embeddings | nomic-embed-text | 768-dim, 8,192-token context window — and critically, the search_document: / search_query: prefix pairing that recovers 5–8% retrieval accuracy most implementations silently discard |
| Vector store | ChromaDB (embedded) | Zero ops overhead — no server process, no port, persists straight to disk |
| Sparse retrieval | BM25Okapi (rank-bm25) |
Catches exact-term matches ("torque specification," "OISD-118 clause 7.4") that pure semantic search misses |
| Score fusion | Reciprocal Rank Fusion, k=60 | Merges dense + sparse rankings without requiring comparable raw score distributions |
| Reranking | ms-marco-MiniLM-L-6-v2 | 22MB cross-encoder — joint query-passage encoding for a precision pass that first-stage retrieval cannot afford |
| Primary database | MongoDB | Schema flexibility for heterogeneous document metadata; native $push / $addToSet / $inc accumulation patterns for the knowledge graph |
| Cache | Memurai (Redis-compatible, Windows-native) | Cache-aside on query results with 1h TTL — repeated demo queries return in single-digit milliseconds |
| Gateway | Node.js + Express | Non-blocking event loop is the right fit for SSE pipe-through and multipart file uploads |
| Frontend | React 18 + Vite + Tailwind + Zustand + D3 | Streaming UI, force-directed knowledge graph, and a design system built to feel less like a dashboard and more like a place you'd actually want to spend eight hours |
elixara/
├── .env.example ← copy to .env before first run
├── package.json ← root orchestration (Windows-adapted)
│
├── frontend/ ← React 18 SPA
│ └── src/
│ ├── pages/ 12 routes: dashboard · upload · query · graph ·
│ │ equipment · failures · compliance · maintenance ·
│ │ analytics · settings · library · login
│ ├── components/ layout · ui primitives · docs · query · graph · charts
│ ├── store/ Zustand slices: app · docs · query · graph
│ ├── hooks/ ingestion polling · SSE streaming · D3 lifecycle
│ └── api/ thin axios wrappers per domain
│
├── gateway/ ← Node.js API Gateway :4000
│ ├── routes/ docs · query · graph · compliance · health · auth
│ └── middleware/ JWT auth · multer upload · morgan logging
│
└── services/ ← Python microservices
├── shared/ config · mongo · chroma · ollama_client · models
├── ingest_service/ :5001 MinerU → chunk → NER → embed → graph build
├── rag_service/ :5002 dense + BM25 + RRF → rerank → stream
├── graph_service/ :5003 node/edge CRUD · BFS pathfinder
└── compliance_service/ :5004 regulation DB · RAG-backed gap scan
Install these once, before touching the project folder. MongoDB, Memurai, and Ollama each register as background Windows services or processes — you will not restart them manually every session.
| Tool | Purpose | Install from | Verify with |
|---|---|---|---|
| Node.js 20+ | Frontend + Gateway | nodejs.org | node --version |
| Python 3.11 | AI microservices | python.org | python --version |
| Git | Version control | git-scm.com | git --version |
| MongoDB Community | Primary database | mongodb.com/try/download/community — .msi, keep Install as a Service checked |
mongosh --version |
| Memurai Developer Edition | Redis-compatible cache | memurai.com/get-memurai | memurai-cli ping → PONG |
| Ollama | Local LLM runtime | ollama.com | ollama list |
Pull both required models once (~2.7GB total):
ollama pull phi4-mini
ollama pull nomic-embed-textOn where you build this. Avoid hosting the project inside a OneDrive-synced folder. OneDrive attempts to sync every file the moment
venvornode_modulescreates it — producing intermittent file-lock errors duringpip install. If you hit unexplained permission errors, move the project toC:\dev\elixaraand the problem disappears.
Every command below assumes you're in the elixara\ project root, in a PowerShell terminal. VS Code's integrated terminal is fine.
git clone <your-repo-url> elixara
cd elixaraCopy-Item .env.example .envThe defaults already point at localhost for MongoDB, Memurai, and Ollama. No edits required for local development.
MinerU carries heavy dependencies (PaddleOCR, ONNX Runtime), so it lives in a dedicated environment outside the project's own venvs:
python -m venv $HOME\.mineru_env
C:\Users\<you>\.mineru_env\Scripts\Activate.ps1Use the full explicit path rather than
$HOMEwhen activating — the$HOME-prefixed form can pick up invisible characters on copy-paste in PowerShell that silently break the command. If PowerShell refuses with "execution of scripts is disabled", run this once and retry:Set-ExecutionPolicy RemoteSigned -Scope CurrentUser
pip install "mineru[core]"MinerU 3.x downloads its parsing models automatically on first real use. Trigger that download now, deliberately, with a real PDF:
mineru -p "C:\path\to\any\test.pdf" -o C:\temp\mineru_test -b pipelineThe first run pulls roughly 1GB of models (layout detection, OCR, table recognition) and can take 5–15 minutes. Every subsequent run is fast. Confirm it worked:
Get-ChildItem -Recurse C:\temp\mineru_test -Filter *.mdOutput structure note. MinerU 3.4 nests output one level deeper than older documentation describes —
<output_dir>\<filename>\auto\<filename>.md, not directly under<filename>\. The codebase already searches recursively withrglob("*.md"), so no code change is needed.
Exit the environment and add MinerU permanently to your PATH:
deactivate
[Environment]::SetEnvironmentVariable("Path", $env:Path + ";$HOME\.mineru_env\Scripts", "User")Fully close and reopen VS Code — not just a new terminal tab. PATH changes only take effect after a full application restart. Then verify:
mineru --versionnpm install
cd frontend && npm install && cd ..
cd gateway && npm install && cd ..MinerU 3.4 pulls in gradio as an optional dependency, which requires newer fastapi and python-multipart than the original pins. Open services\shared\requirements.txt and confirm:
fastapi>=0.115.2
python-multipart>=0.0.18
If they still show exact pins, loosen them — otherwise pip install will fail with a ResolutionImpossible error.
ingest_service:
cd services\ingest_service
python -m venv .venv
.\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..rag_service — install CPU-only PyTorch before the requirements file, or sentence-transformers will pull a CUDA build:
cd services\rag_service
python -m venv .venv
.\.venv\Scripts\pip.exe install torch==2.3.1+cpu --index-url https://download.pytorch.org/whl/cpu
.\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..graph_service and compliance_service:
cd services\graph_service
python -m venv .venv && .\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..
cd services\compliance_service
python -m venv .venv && .\.venv\Scripts\pip.exe install -r requirements.txt
cd ..\..mongosh --eval "db.runCommand({ ping: 1 })"
memurai-cli ping
ollama listIf Ollama shows no models, pull phi4-mini and nomic-embed-text as listed in Prerequisites. If Memurai doesn't respond, open services.msc and confirm the Memurai service is set to Running.
From the project root, one command starts everything — frontend, gateway, and all four Python microservices, concurrently:
npm run dev:allGive it a few seconds. You're looking for six clean startup lines with no traceback beneath any of them:
[dev:frontend] VITE ready
[dev:gateway] ⬡ Elixara Gateway running on :4000
[dev:ingest] Uvicorn running on http://127.0.0.1:5001
[dev:rag] Uvicorn running on http://127.0.0.1:5002
[dev:graph] Uvicorn running on http://127.0.0.1:5003
[dev:compliance] Uvicorn running on http://127.0.0.1:5004
Then open:
Demo login: demo / elixara2024 — or hit ⚡ Judge Demo Access for an instant token.
python scripts\seed_demo.pyUploads sample PDFs through the full ingestion pipeline. If demo_docs\ doesn't exist yet, add your own PDFs first — or drag files directly onto the Upload page in the running app.
python scripts\check_health.pyOr open the Settings page inside the app (http://localhost:5173/settings) — it live-polls every service and Ollama itself, showing a green or red dot with latency for each.
A fully healthy stack:
{
"ingest": { "status": "ok" },
"rag": { "status": "ok" },
"graph": { "status": "ok" },
"compliance": { "status": "ok" },
"ollama": { "status": "ok" }
}| Service | Framework | Port | Interactive docs |
|---|---|---|---|
| Frontend | React 18 / Vite | 5173 | http://localhost:5173 |
| API Gateway | Express | 4000 | http://localhost:4000/api/* |
| Ingest Service | FastAPI | 5001 | http://localhost:5001/docs |
| RAG Service | FastAPI | 5002 | http://localhost:5002/docs |
| Graph Service | FastAPI | 5003 | http://localhost:5003/docs |
| Compliance Service | FastAPI | 5004 | http://localhost:5004/docs |
| Ollama | — | 11434 | http://localhost:11434 |
| MongoDB | — | 27017 | Windows Service |
| Memurai | — | 6379 | Windows Service |
| Symptom | Root cause | Fix |
|---|---|---|
mongosh not recognized |
MongoDB not installed, or terminal opened before install | Reinstall via .msi with Install as a Service checked; open a fresh terminal |
memurai-cli not recognized |
Same, for Memurai | Reinstall; confirm the Memurai service is Running in services.msc |
mineru not recognized after setting PATH |
VS Code caches environment variables per-window | Fully close all VS Code windows and reopen — a new terminal tab alone is not enough |
PowerShell parser error on $HOME\.mineru_env\... |
Invisible characters from copy-paste | Retype manually or use the full explicit path instead of $HOME |
ResolutionImpossible on fastapi or python-multipart |
MinerU's gradio dependency requires newer versions than the original pins |
In services/shared/requirements.txt: loosen to fastapi>=0.115.2 and python-multipart>=0.0.18 |
TypeError: unsupported operand type(s) for | in chroma.py |
chromadb.PersistentClient is a factory, not a class — the X | None type union fails at import time |
Add from __future__ import annotations as the first import in services/shared/chroma.py |
NameError: name 'logging' is not defined in scan.py |
Missing standard-library import | Add import logging at the top of compliance_service/routers/scan.py |
"The system cannot find the path specified" for dev:graph or dev:compliance |
That service's .venv was never created |
Run Step 5 for the missing service |
.venv/bin/uvicorn errors |
Copy-pasted Linux path syntax | Windows uses .venv\Scripts\uvicorn, not .venv/bin/uvicorn |
pip install fails on torch |
Wrong index URL or missing CPU flag | Use exactly: pip install torch==2.3.1+cpu --index-url https://download.pytorch.org/whl/cpu |
| Port already in use | A previous process didn't shut down cleanly | Get-Process -Id (Get-NetTCPConnection -LocalPort 5001).OwningProcess | Stop-Process |
| SSE answers appear all at once instead of token-by-token | Proxy buffering | Confirm res.flushHeaders() runs and X-Accel-Buffering: no is set — already configured in this repo's gateway |
| ChromaDB "collection not found" | Corrupted local vector store | Delete data\chroma_db\, restart the RAG service, re-ingest |
| Symlink warning during MinerU model download | Windows blocks symlinks without Developer Mode | Harmless — Hugging Face falls back to full file copies; no functional impact |
| Intermittent file-permission errors during dependency install | Project inside a OneDrive-synced folder | Move to C:\dev\elixara |
There is something almost stubborn about the idea at the center of this project.
Not that the problem is hard — though it is. Not that the technology is novel — though it is assembled carefully. But that the answer was already there. Not behind a paywall. Not requiring a new sensor, a new survey, a new budget line. Buried, instead, under the ordinary weight of twelve disconnected systems and a PDF nobody has opened since the inspection that produced it.
I built Elixara because I kept thinking about those retiring engineers — the ones who carry in their heads the reason a particular pump keeps failing, the way a particular regulation intersects with a particular equipment class, the shortcut that saved three hours the last time this exact alarm fired. When they leave, that knowledge doesn't get archived. It evaporates.
Elixara doesn't claim to know more than your engineers do. It claims to remember what they already knew, and to say it back the moment someone needs it — with the receipt attached, every time, so that trust is never asked for without being earned.
The plant already has a brain. Elixara is how it finally speaks.
Distributed under the MIT License. See LICENSE for details.
⬡ Elixara · CPU-Only · Zero Cloud Cost · Zero Hallucination Tolerance
Built for the industrial engineers of tomorrow — and the knowledge they carry today.