Multi-agent AI copilot for data engineers — LangGraph orchestration, specialist agents, Databricks tools, RAG, and MLflow.
- LangGraph —
rag → classify → pyspark | sql | airflow | workflow → validation - CDC workflow — plan then parallel SQL + PySpark (docs)
- Human approval — SQL execution gate +
POST /approval/{id}(docs) - Guardrails — input/SQL/output validation
- Agent memory — per-
session_idconversation history - Retry logic — Ollama + tools with backoff
- Notebook export — optional JSON export path
- Keyword tools — SQL, Delta, schema, Vector Search, MLflow (docs)
- RAG — Vector Search context on every chat
- Local LLM — Ollama (
llama3.2default) - MLflow — one run per chat (optional Databricks creds)
- Notebooks — Databricks setup 00–04
- Ollama · FastAPI · Streamlit · LangGraph
- Databricks Vector Search · Spark SQL · MLflow
docs/README.md — full index.
cp .env.example .env
make start| Service | URL |
|---|---|
| Streamlit | http://127.0.0.1:8501 |
| API | http://127.0.0.1:8000 |
| Swagger | http://127.0.0.1:8000/docs |
CDC workflow:
curl -X POST http://127.0.0.1:8000/chat \
-H "Content-Type: application/json" \
-d '{"question":"design a cdc pipeline for customer orders"}'{
"agent_used": "workflow_agent",
"graph_branch": "workflow",
"tools_used": ["planner", "sql_agent", "spark_agent"],
"rag": { "context_size": 90, "error": null }
}Schema / SQL branch:
curl -X POST http://127.0.0.1:8000/chat \
-H "Content-Type: application/json" \
-d '{"question":"schema of main.ai_copilot.rag_documents"}'See .env.example — LLM_MODEL, DATABRICKS_HOST, DATABRICKS_TOKEN, optional MLFLOW_EXPERIMENT.
pytest tests/ -v
./scripts/test_workflow.sh