All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
-
Model-initiated web search. When enabled, the model can call two tools on its own:
web_search(DuckDuckGo — free, no API key or setup) andfetch_page, which retrieves a page and returns its main content as Markdown with headings and fenced code blocks preserved, so API documentation stays readable. The model decides when to search, reads what it finds, and answers with citations; every query and fetched URL is listed in a collapsible "Searched the web" block under the message, and links open in the browser. -
Web search is off by default — queries and URLs leave your machine — and can be toggled per session with the 🌐 button in the chat toolbar. Configure via
insightcoder.webSearch.*:enabled,maxResults(5),maxHops(5, each hop re-sends the project context),fetchPages,maxPageChars(8000) andregion. -
Page fetching is hardened against the model choosing hostile URLs: http(s) only, loopback/private/CGNAT ranges and the cloud metadata endpoint are refused, redirects are capped and re-validated at every hop, responses are size-capped and content-type checked, and fetched text is delivered to the model explicitly labelled as untrusted data rather than instructions.
-
DeepSeek support. Select
deepseekas the provider and usedeepseek-reasoner; its chain of thought appears in the same collapsible thinking blocks as other reasoning models. DeepSeek keeps its own API key alongside MiniMax and Gemini, so you can switch between all three mid-conversation. -
The model dropdown lists every provider, including ones you haven't configured — entries for a provider with no stored key are marked "· key needed", and selecting one prompts you to add the key. Only reasoning-capable models are offered (whole-codebase analysis is a reasoning task), and superseded, preview, dated and non-chat variants are filtered out, so a provider's live model list can't flood the dropdown.
- Each provider now resolves its own endpoint.
insightcoder.baseUrlchanged from a fixed MiniMax default to an empty override: leave it blank and every provider uses its documented endpoint (MiniMaxhttps://api.minimax.io/v1, DeepSeekhttps://api.deepseek.com/v1); set it and it applies only to the provider selected in settings, so switching providers in the dropdown can no longer send requests to another provider's URL. - A reply is now shown as the ordered timeline of what actually happened — think, search, write, think again — instead of one fixed reasoning block and one sources block above the text. Each thinking block holds only the thoughts from that step, each search appears where it occurred, and everything stays expandable after the turn ends. Conversations saved before this update are rendered through the same path.
- A message could end with no answer at all: if the model was still requesting tools when the hop budget ran out, the loop exited without producing text. The final hop is now run without tools, so the model always has to answer.
- Text the model wrote before calling a tool was discarded when the next hop began, making a completed answer visibly disappear. It is kept, and the post-tool continuation appends to it.
- The "searching" indicator only appeared after the response finished, because tool calls were surfaced at the end of the stream. Tools are now announced as soon as they are recognised mid-stream.
- Reasoning traces could be overwritten rather than accumulated across tool round-trips, so earlier thinking was lost.
- Message editing with conversation branches. Any message can be edited via a pencil button that appears on hover. Editing a user message creates a new branch and immediately resends it; editing a model response creates a branch with the corrected text. The original branch is always preserved — a ChatGPT-style ‹ 2/2 › switcher appears on messages that have alternative versions, and switching restores that branch's own downstream conversation. Conversations are now stored as trees (existing linear conversations are migrated automatically), and only the active branch is sent to the model, summarized, and exported.
- Toolbar split into two lines. The file count and input-token readout moved to their own row below the model selector and action buttons, so the information stays readable in a narrow sidebar. The model dropdown now stretches to fill the freed space.
- MiniMax stopped thinking on follow-up questions. MiniMax M-series interleaved thinking requires assistant turns in the conversation history to be sent back with their
<think>reasoning blocks; the extension was stripping them (correct for other providers), which caused reasoning to degrade and disappear after the first turn. Stored thinking traces are now restored into history for the MiniMax/OpenAI-compatible provider only (Gemini still receives clean history), and the token estimate accounts for the replayed reasoning. - "ConversationStore not initialized" on startup. The webview's ready signal could win the race against the asynchronous conversation-store initialization on a fresh activation, surfacing an error banner and preventing the context (and the file/token readout) from loading. Webview messages now wait for store initialization to complete.
- Panel restore is now instant. Switching away from the InsightCoder tab and back previously took several seconds: VS Code destroyed the hidden webview, and re-initialization blocked on two sequential network calls fetching provider model lists. The webview is now retained while hidden, initialization uses a cached/static model list with zero network I/O (the dropdown refreshes in the background, with both providers queried in parallel), and the last conversation is snapshotted so even a genuinely reloaded webview (window restart) repaints immediately.
- Context inspector. A new
InsightCoder: Inspect Context Filescommand opens a searchable, multi-select list of every file in the current model context, sorted largest-first and annotated with size (KB), estimated tokens, and each file's share (%) of the context. Ticking files and confirming excludes them, their paths are appended toinsightcoder.context.excludein the workspace settings and the context is rebuilt immediately (Explorer badges and the token counter update). Launchable from the Command Palette, a new 📊 button in the chat toolbar, and a "Manage large files…" button on the over-limit warning banner.
- Reasoning traces. When a model exposes its thinking, it is shown in a collapsible block that is hidden by default. While the model is thinking, an animated indicator makes the process clearly visible and a continuously updating count shows how many tokens are being spent on reasoning; the completed trace stays available to expand. Supports Gemini thoughts and MiniMax M-series inline
<think>…</think>reasoning. - Resend unanswered messages. If a message fails to reach the model (network error, missing API key, etc.), the unanswered message can be resent with one click instead of being retyped.
- Context visibility in the Explorer. Files included in the current model context are marked with a check badge; files that were scanned but skipped get a minus badge with the reason. Unmarked files are not part of the context.
- Open the full assembled context in an editor tab from a toolbar button in the chat panel (in addition to the existing
InsightCoder: Show Assembled Contextcommand). - Upfront input token count. The panel shows how many input tokens the next message will cost before anything is sent (exact on Gemini, estimated on MiniMax), updating live as you type and after each turn.
- Default Gemini models updated to
gemini-3.5-flashandgemini-3.5-pro. - Summaries exclude reasoning traces — thinking is stripped from generated summaries so it is never stored as long-term memory.
- MiniMax M-series reasoning was rendered as plain answer text with no way to hide it; its inline
<think>…</think>reasoning is now parsed out of the response stream (handling tags split across chunks) and routed to the collapsible thinking block. - Reasoning no longer disappeared on follow-up turns: continuation responses whose chat template auto-prefills the opening
<think>(so only a closing</think>is streamed) are now detected and routed to the thinking block instead of leaking into the answer.
- Complete rewrite as a VS Code extension (TypeScript). The standalone Python/Flet desktop app is retired; InsightCoder is now installable as a
.vsix/ from the Marketplace. - Conversations moved out of the repo. Chats are stored as JSON in the extension's own global storage (per workspace) with a conversation-tab UI — nothing is written into the analyzed project anymore.
- Context engine rebuilt: one
git ls-filescall instead of per-filegit check-ignoresubprocesses, binary detection by content sniffing instead of an extension allowlist, include/exclude globs in settings, per-file size cap, and an explicit confirmation flow when the estimated context exceedsinsightcoder.context.maxTokens. - The
full_context.txtdebug dump is replaced by theInsightCoder: Show Assembled Contextcommand.
- Multi-provider support: MiniMax M-series (or any OpenAI-compatible endpoint via
insightcoder.baseUrl) as the primary provider, plus Google Gemini — switchable from a model dropdown mid-conversation. - Secure API key storage in VS Code Secret Storage via
InsightCoder: Set API Key…(replaces theGEMINI_API_KEYenvironment variable). - Conversation tabs (new/switch/delete) and
Export Conversation as Markdown. - Live token counter: exact counts on Gemini; self-calibrating estimates on MiniMax (recalibrated from real usage each turn).
- Streaming cancellation (Stop button), chunk batching, and sanitized Markdown rendering (DOMPurify + highlight.js).
- Unit test suite (vitest) for the context engine, prompt builder, conversation store, and token service;
npm run test:livesmoke test against real provider endpoints.
- Model Selector: Implemented a dropdown menu in the UI, allowing users to choose between
gemini-2.5-proandgemini-2.5-flashmodels for their conversation. - Selectable Chat Text: All chat messages, including the welcome screen, are now selectable, making it easy to copy code snippets and responses.
- Enhanced Welcome Message: The initial welcome screen now includes a more detailed message with a privacy notice and clearer getting-started instructions.
- Project Structure Refactoring: Reorganized the project's file structure to improve clarity and maintainability.
- The legacy
ask_srcdirectory has been removed. - All backend modules (
services.py,chat_utils.py) have been consolidated into a new, cleanersrcdirectory. - The main application entry point was renamed from
flet_ask.pytoask.py, simplifying the project's root. - All internal imports were updated to reflect the new file paths.
- The legacy
- UI Responsiveness: Resolved an issue where the UI would freeze for a moment when a message was sent. The application now yields to the UI thread to render the user's message before making the blocking API call.
- Complete Frontend and Architectural Rewrite: The entire application frontend has been migrated from PyQt5 to Flet. This is a ground-up rewrite aimed at modernizing the codebase, improving maintainability, and enabling simple cross-platform builds.
- Modernized Concurrency Model: Replaced the previous multi-threaded worker architecture (
ChatWorker,TokenCountWorker, etc.) and PyQt's signal/slot system with a nativeasyncioimplementation. This simplifies the code, eliminates complex threading logic, and improves efficiency.- Blocking operations (like API calls and file I/O) are now handled non-blockingly using
asyncio.to_thread. - Real-time features like the token counter's debouncing are now elegantly handled with
asyncio.sleep, removing the need forQTimer.
- Blocking operations (like API calls and file I/O) are now handled non-blockingly using
- New Code Architecture: Introduced a clean, three-tier architecture to separate concerns:
- State Management (
AppState): A central class holds all application state, making data flow predictable. - Service Layer (
app_services.py): All backend logic (API calls, file operations, context management) is encapsulated in asynchronous services (ChatService,ContextService, etc.). - UI Components: The UI is broken down into logical Flet components (
ChatView,InputBar) for better organization and reusability.
- State Management (
- UI Rendering: Replaced
QWebEngineViewwith Flet's nativeft.Markdowncontrol for robust, built-in rendering of chat messages and code blocks.
- PyQt5 and PyQtWebEngine Dependencies: Removed all PyQt-related libraries from
requirements.txt, significantly reducing the project's dependency footprint. - Obsolete UI and Worker Code: Deleted the legacy PyQt UI file (
ask_src/ui.py), the original entry point (ask.py), and the PyQt-specific worker patterns. The new application entry point is the refactoredflet_ask.py.
- Syntax Highlighting Engine: Replaced the server-side
Pygmentslibrary with the client-sidehighlight.jslibrary for rendering syntax-highlighted code blocks. This change leveragesQWebEngineViewto delegate highlighting to JavaScript, resulting in significantly more robust and accurate code formatting in the chat display. - UI Rendering Logic: Refactored
ask_src/ui.pyandask_src/worker.pyto remove thecodehiliteMarkdown extension and instead loadhighlight.jsassets and trigger the highlighting script within theQWebEngineView.
PygmentsDependency: RemovedPygmentsfromrequirements.txtand deleted the obsoletepygments_default.cssfile, simplifying the project's dependencies.
- "Reload Context" Button: Introduced a "Reload Context" button in the UI. This allows users to re-scan the project's codebase on demand, updating the AI's knowledge with the latest file changes without losing the current conversation history.
ContextReloadWorker: Implemented a new background worker inask_src/worker.pyto handle the context reloading process asynchronously. This ensures the UI remains responsive while a new chat session is prepared with the updated context and existing conversation history.- Welcome Screen: Added an initial welcome message that is displayed on application startup. It provides guidance and example prompts for new users and is replaced by the conversation once the first message is sent.
- Chat Rendering Engine: Replaced the
QTextBrowserwidget withQWebEngineViewfor displaying chat messages. This change significantly improves the rendering quality and accuracy of Markdown and syntax-highlighted code blocks by leveraging a full web engine. - UI Logic: Refactored
ask_src/ui.pyto supportQWebEngineView, which involved embedding styled content within a full HTML structure and handling UI interactions like scrolling via JavaScript.
- Conversation Summarization: Implemented an intelligent system to automatically summarize conversations using the LLM. This provides long-term memory while significantly reducing the token count sent to the model, making context management more efficient and cost-effective.
SummaryWorker: Introduced a new background thread inask_src/worker.pyto handle conversation summarization asynchronously. This ensures the UI remains responsive while summaries are being generated after a conversation is saved.- Startup Summarization: Added logic to
ask.pythat automatically detects and summarizes any previously unsummarized conversation files when the application starts. This ensures all history is processed into efficient summaries before the chat session begins. .gitignoreIntegration: Theget_codebasefunction now respects.gitignorerules by using thegit check-ignorecommand, leading to a more accurate and relevant codebase context.reconstruct_markdown_history: A new helper function inchat_utils.pyto convert saved conversation data back into a full markdown string, which is necessary for the summarization process.
- Context Management: The system prompt (
create_system_prompt) now loads concise conversation summaries instead of full conversation histories, optimizing token usage and improving performance. - Application Startup Flow: Refactored
ask.pyto configure the API client once at startup and pass it to all necessary components. The startup sequence now completes all pending summarization tasks before launching the main UI. - Conversation Saving: The
save_conversation_mdmethod inui.pynow triggers theSummaryWorkerto create a summary file after successfully saving the full conversation markdown.
- Conversation Counter: Corrected an issue where the conversation counter could be miscalculated. It now accurately determines the next conversation number by checking for both
conversation_N.mdandconversation_N_summary.mdfiles. - Path Exclusion: Improved the logic for excluding the conversation directory from codebase analysis, making it more robust, especially when custom conversation paths are used.
- Refined
README.mdfor greater conciseness and accuracy, updating information on features and getting started. - Updated the referenced LLM model name (
gemini-2.5-flash-preview-04-17) and its context window size (up to 1M tokens) inREADME.mdto reflect the model used inchat_utils.py. - Updated main prompt for better context and clarity.
- Resolved an issue when the conversation were not saved properly after the LLM complete the output.
- Command-line option
--conversation-path(or-c) toask.pyfor specifying a custom directory for conversation history, providing users with control over conversation storage location. .gitignorehandling for conversation folder: Implemented logic to ensure conversation loading works correctly even when the conversation folder (default or custom) is ignored by Git, preserving conversation history while respecting.gitignorefor codebase analysis.- Diff Detection in AI Responses: Implemented logic in
diff_detector.pyto automatically detect code diff blocks within the AI's responses, identifying potential code modifications suggested by the AI. - File Path Extraction from Diff Blocks: Enhanced diff detection to extract file paths from the headers of detected diff blocks, enabling identification of target files for code changes.
- File Path Validation: Implemented validation logic to check if extracted file paths from diff blocks correspond to existing files within the analyzed project, preventing application of diffs to invalid or non-existent files.
- Unit Tests for Diff Detection: Added comprehensive unit tests in
tests/test_diff_detector.pyfor thedetect_diff_blocksfunction, ensuring the robustness and correctness of diff detection and file path extraction logic. - Test for Conversation Folder Handling: Added integration tests in
tests/test_conversation_folder.pyto verify correct conversation folder handling, including.gitignorescenarios and custom conversation paths.
- Refactored
detect_diff_blocksfunction indiff_detector.pyto return a list of dictionaries, each containing diff block content and extracted file path, improving the structure of diff block information. - Updated
ChatWorkerinworker.pyto utilize the newdetect_diff_blocksfunction for processing AI responses and logging detected diff blocks with file path information to the console. - Modified
ask.pyandui.pyto accept and pass theconversation_pathcommand-line argument, enabling users to specify custom conversation folders. - Updated
ROADMAP.mdto reflect the implemented features and highlight next priorities, providing an up-to-date project development plan.
- Implemented background token counting using
TokenCountWorkerto prevent UI freezes during text input, ensuring a smoother user experience. - Added
TokenCountSignalsfor signal-based communication between theTokenCountWorkerthread and theMainWindowUI thread, enabling thread-safe updates of the token count display. - Introduced a visual "Tokens: [count]" label in the UI, positioned above the input text area, to display real-time token counts to the user.
- Implemented debouncing for token count updates using
QTimer, reducing the frequency of token counting calculations and further improving UI responsiveness, especially during rapid typing. - Added a
ROADMAP.mdfile to the project root, outlining the project's short-term, medium-term, and long-term development vision and planned features. - Included an MIT License to the project, adding a
LICENSEfile in the root directory and a "License" section toREADME.md, clarifying the open-source licensing terms. - Added a comparison table with GitHub Copilot to
README.md, highlighting the key differentiators and complementary nature of InsightCoder. - Added a "Beyond Code: Analyzing Any Git Repository" section to
README.md, emphasizing that InsightCoder can analyze any Git repository, including documentation, books, and other non-code projects.
- Refactored token counting logic: Moved the token counting functionality from the
update_token_count_displaymethod inMainWindow(ui.py) into the dedicatedTokenCountWorkerclass (token_worker.py), promoting code modularity and separation of concerns. - Updated
ui.pyto utilizeTokenCountWorkerfor asynchronous token counting, ensuring non-blocking UI operations. - Modified
ui.pyto include new methodsstart_token_count_timerandset_token_count_labelfor managing the debounced token counting and updating the UI label via signals and slots. - Improved UI responsiveness and smoothness during text input and token count updates by offloading the potentially time-consuming token counting process to a background thread, preventing UI freezes.
- Streamlined the "Contributing" section in
README.mdto be more concise and focused on bug reports, aligning with the project's current self-development stage.
- Resolved UI freezing issue that occurred during text input due to synchronous token counting in the main UI thread, significantly enhancing the user experience.
- Command-line argument
--project-path(or-p) toask.pyfor specifying the project directory to analyze. - Documentation in
README.mdon how to use the--project-pathargument.
- Window title dynamically updates to display the name of the project directory being analyzed (e.g., "[Project Name] Codebase Chat").
- System prompt for the LLM is now more generic, focusing on codebase analysis rather than being specific to "InsightCoder" project identity.
MainWindowinui.pynow acceptsproject_pathas an argument for cleaner code and dynamic title setting.- Updated
ask.pyandchat_utils.pyto handle and pass theproject_pathargument correctly.
- Addressed the issue of the project analyzing only its own codebase by implementing
--project-pathfunctionality. - Resolved the confusion of "InsightCoder" project identity when analyzing external codebases by making the system prompt and window title more context-aware.
This is the initial release of InsightCoder as a standalone open-source project, extracted from the Anagnorisis project.
- Initial codebase for InsightCoder, providing AI-powered codebase analysis.
- Interactive chat interface using PyQt5.
- Integration with Google Gemini API for natural language code queries.
- Markdown formatted output with syntax highlighting in responses.
requirements.txtfile for easy installation of dependencies.- Basic
README.mddocumentation to get users started. - Saving conversation history to
.mdfiles.
- Project renamed from internal tool to InsightCoder for open-source release.
- System prompt and UI updated to reflect the standalone InsightCoder identity.
- Conversation files are now saved in the
project_info/conversationsdirectory.
- Grooveshark related notes from
README.md. - Anagnorisis specific branding and references throughout the codebase.