Skip to content

Repository files navigation

market//recon

Source-backed market sizing in minutes, not weeks — an open-source LLM operator that researches public data and shows every receipt.

Dashboard — TAM/SAM/SOM with live agent pipeline

The problem

Early-stage teams and GTM/strategy functions constantly need a defensible answer to "how big is this market?" — for a fundraise deck, a board memo, a product bet. The options today are bad: analyst reports cost thousands and go stale, DIY spreadsheet sizing has no sources behind it, and a single ChatGPT prompt produces confident numbers that cannot be audited. What these teams actually need is fast, source-backed market sizing where every figure traces to a public source and every assumption is visible.

The solution

market-recon is a self-hosted middleware: plug in basic product info (name, description, target customer, pricing model, geography) and an open-source LLM acting as an autonomous operator researches the market with OSINT tools — web search, public-page scraping, open datasets — then returns a structured TAM/SAM/SOM analysis with segmentation, a competitive landscape map, and a complete citation trail.

Three properties make the output credible rather than confident:

  1. Every number shows its work — formula, inputs, sources, and a confidence grade, rendered in a "sources" drawer next to each figure.
  2. Honesty about gaps — anything the agent could not verify from at least one public source is stamped NO DATA FOUND, never invented. The synthesizer is structurally prevented from citing URLs it was not actually shown.
  3. Nothing proprietary goes in, nothing is retained — it queries public sources at request time and keeps results in memory only.

Why an agent, not a prompt

A single prompt call asks the model to remember the market. This pipeline makes the model research it — and confines each LLM judgement to a step where it can be checked:

Step Role LLM? Failure containment
1 Query Planner — decompose product into research sub-questions yes falls back to a template plan
2 OSINT Researcher — search, fetch, build the evidence registry no per-query retry; a dead backend becomes a gap, not a crash
3 Data Synthesizer — extract figures/competitors/pricing with citations yes facts citing unknown URLs are dropped and logged
4 Market Sizing Analyst — apply the TAM/SAM/SOM formulas no pure functions; the LLM never does arithmetic
5 Report Composer — segmentation + positioning map yes deterministic fallback, degradation stated in warnings

After synthesis the operator reflects: empty evidence categories trigger one bounded re-search pass with alternative queries. The loop is a ~120-line custom plan-execute-reflect module (marketrecon/agent/operator.py) rather than a framework graph — every transition is visible, and the whole pipeline runs offline in tests against a fake LLM and a recorded mini-web.

The model is swappable via env alone — any OpenAI-compatible endpoint works (Ollama, vLLM, LM Studio, llama.cpp-server): set LLM_BASE_URL and LLM_MODEL. Structured output uses schema-in-prompt JSON with a parse → fence-strip → bracket-slice → repair-prompt ladder, which is far more robust across open models than native tool-calling.

Quickstart

Demo mode — zero dependencies. Replays a canned analysis of the fictional company Beacondesk through the real SSE pipeline, so you can explore the full animated dashboard offline:

uv sync
cd frontend && npm install && npm run build && cd ..
uv run uvicorn marketrecon.api.app:app --port 8000
# open http://localhost:8000 and click “Replay demo analysis”

(No uv? Plain pip works too: pip install -e . and run uvicorn marketrecon.api.app:app.)

Live mode. Point it at an open LLM and a search backend:

cp .env.example .env        # defaults: Ollama at localhost:11434, SearxNG at localhost:8080
ollama pull llama3.1        # the model must actually be pulled, or you'll get a clear 404 hint
docker compose up           # api + SearxNG (add --profile local-llm for bundled Ollama)

.env is honored by direct uvicorn runs as well as docker-compose. If the LLM or search backend is missing, the run fails fast with a human-readable message on the dashboard (and demo mode always works with neither).

Everything in Docker:

docker compose --profile local-llm up -d
docker compose exec ollama ollama pull llama3.1
LLM_BASE_URL=http://ollama:11434/v1 docker compose up api

Methodology

All three standard sizing methods run per analysis (selectable in the intake form), and the output labels each one:

Top-down — anchors on analyst/industry-report figures the researcher finds, aligned to the product category:

TAM = median(industry report market values)

Bottom-up — built from OSINT-discovered customer counts and pricing benchmarks:

TAM = potential_customers × avg_revenue_per_customer
      (ARPU = user-provided price point, else median public benchmark)

Value-theory — from what comparable products charge and the premium/discount the product's positioning implies:

TAM = comparable_price × positioning_factor × potential_customers
      (factor = own_price ÷ comparable_median, clamped to [0.5, 2.0]; 1.0 if unknown)

SAM and SOM apply the standard filters on top of each method's TAM:

SAM = TAM × serviceable_share     (geographic / demographic / product-fit constraints,
                                   estimated by the synthesizer from evidence; defaults
                                   to 30% with the default explicitly flagged)
SOM = SAM × obtainable_share

obtainable_share is a documented heuristic (marketrecon/sizing/methods.py), not model vibes: a company-stage base (pre-launch 2% / early customers 5% / scaling 10%) damped by competitor density (halved at 10 competitors), scaled by market saturation (low ×1.25 / medium ×1.0 / high ×0.6), clamped to [0.5%, 25%]. The full derivation string is included in every report's assumptions.

Confidence grades (marketrecon/sizing/confidence.py) are equally auditable:

  • high — corroborated by ≥ 2 independent domains, all retrieved within 180 days
  • medium — a single independent source, or corroborated but stale
  • low — no public source (the value will also be flagged unverified)

Medians (never single cherry-picked values) aggregate multiple discovered figures, which also makes the pipeline robust to one bad extraction.

Segmentation and competitive landscape

Segmentation & competitive landscape

  • Segmentation across three axes — firmographic, geographic, use-case — each segment sized as share × SAM with the share's provenance noted (LLM-inferred shares are capped at medium confidence).
  • The researcher identifies direct and adjacent competitors; the synthesizer extracts positioning, pricing tier, and market-share signals (funding, headcount, review volume) purely from public sources.
  • The composer chooses the two most informative axes and places competitors on a positioning quadrant the frontend renders as an animated scatter — hover any dot for its sourced details.

Architecture

frontend/   Vite + vanilla JS + anime.js v4 dashboard (SSE progress, animated report)
api/        FastAPI: POST /api/analyses · GET …/events (SSE) · …/report · …/sources
agent/      operator loop + 5 roles — HTTP-agnostic, unit-testable
tools/      SearxNG / SerpAPI search · cached fetcher (SQLite, TTL) · World Bank data
sizing/     pure TAM/SAM/SOM math + confidence scoring
schemas/    versioned Pydantic contract (Report v1.0) shared with the frontend

Design notes live in docs/DESIGN.md. The UI respects prefers-reduced-motion, imports only the anime.js modules it uses (~20 KB gzip total JS), and its chart palette is validated for color-vision deficiency and contrast on both light and dark surfaces.

Testing

uv run pytest        # 34 tests, fully offline

The suite covers exact sizing math, the JSON-repair ladder, tool adapters against recorded responses, the full agent pipeline against a fake LLM (event ordering, citation integrity, gap honesty, hallucinated-source rejection, composer fallback), and the API contract end-to-end over demo mode.

Configuration

Variable Default Purpose
LLM_BASE_URL http://localhost:11434/v1 any OpenAI-compatible endpoint
LLM_MODEL llama3.1 e.g. qwen3, glm-4.5-air, mistral
LLM_API_KEY ollama placeholder locally; real key for hosted endpoints
SEARCH_BACKEND searxng searxng (keyless) or serpapi
SEARXNG_URL http://localhost:8080 self-hosted instance
SERPAPI_KEY — only if using SerpAPI
CACHE_PATH / CACHE_TTL_HOURS data/cache.sqlite3 / 24 scrape cache

Data & privacy

No proprietary or customer data is required, requested, or stored. Analyses run against public OSINT sources fetched at request time; runs live in process memory and disappear on restart. The only persistence is a local TTL-bound cache of already-public pages — delete the file to forget everything. The examples/ folder contains a clearly fictional company with *.example source domains.

What this demonstrates

  • Agent orchestration — a plan-execute-reflect operator with role isolation, bounded retries, and graceful degradation, testable end-to-end offline.
  • OSINT tool integration — swappable search backends, polite cached scraping, keyless open-data enrichment, and citation tracking from raw fetch to final figure.
  • Market-sizing rigor — three named methodologies with explicit formulas, a documented obtainable-share heuristic, confidence grading, and structural honesty about missing data.
  • Frontend engineering — a live agent-progress stream and an animated, accessibility-conscious report UI over a versioned JSON contract.

Framed around the business outcome: a founder gets a defensible, auditable market size — with sources — from a laptop, for free.

Pairs with revops-forecast (CRM-agnostic revenue forecasting) — same visual language, same "show your work" ethos.

License

MIT

About

OSINT market-sizing middleware — an open-source LLM operator that researches public data and outputs auditable TAM/SAM/SOM analyses with full citation trails

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages