This project is a Spring Boot chatbot that lets users upload their own files and ask questions about them, using Retrieval-Augmented Generation (RAG) with a local Ollama LLM.
Typical flow:
- Login — Access is protected by Spring Security form login (
/login); each request runs under an authenticated user. - Simple UI — A Thymeleaf-based chat page (
home.html) lets a logged-in user upload files and chat with the assistant directly in the browser.
From the user's perspective, that's it — log in, upload a file, and start chatting. Under the hood, here's what actually happens on each of those actions:
- Upload a document —
POST /api/uploadaccepts a file (PDF, DOCX, TXT, etc., up to 50MB). The file is parsed with Apache Tika, split into chunks, embedded (via thenomic-embed-textOllama model), and stored in a PGVector vector store. - Chat about the uploaded content —
POST /api/ai/chatsends the user's question to aChatClientbacked by thegemma4Ollama model. AQuestionAnswerAdvisorretrieves the most relevant chunks from the vector store (top 6 matches, similarity threshold 0.3) and injects them into the prompt, so answers are grounded in the user's own uploaded documents rather than the model's general knowledge. - Conversation memory — A
MessageChatMemoryAdvisorkeeps a per-user rolling chat history (keyed by the authenticated principal), so follow-up questions retain context.
In short: it's a self-hosted chat with your documents assistant — useful for querying manuals, reports, or notes without sending data to a third-party AI service, since both the LLM and the vector store run locally.
-
Install Ollama in development environment and pull
gemma4model. -
Start PostgreSQL Server with PgVector, Use provided docker compose file.
docker compose -f pgvector-docker-compose.yaml up -d
gradle clean bootRun
- Stop Application
Press Ctrl+C to stop the application. Stop the PostgreSQL database using following command.
docker compose -f pgvector-docker-compose.yaml down -v