Ask leading AI models to rank Italian life insurers from a customer's point of view, and see what each model says about each company.
Customers used to compare life insurance through agents, bank advisers, comparison sites and word of mouth. More and more of them now start by asking an AI assistant such as ChatGPT, Claude or Gemini:
"We're a young family in Italy looking for our first life insurance. Which company should we go with?"
The answer they get shapes their shortlist before they ever visit your website or speak to an adviser. Yet most insurers have no idea what these models say about them:
- Where does your company appear in the ranking: near the top, in the middle, or last?
- Why does the model place you there? Brand strength, financial solidity, product complexity, costs, distribution?
- Do different models agree with each other, and does the same model give the same answer twice?
- Is the model's picture of your company accurate, or out of date?
This repository makes those questions measurable. It puts the same question, framed as a specific customer, to several AI models and turns their answers into comparable tables of ranks and reasons, one row per company.
companies.csv ─┐
├─► prompt ─► each model (× runs) ─► parse table ─► results/<timestamp>/<persona>/
personas/*.txt ┘
For each persona in personas/, and for each model, and for each run:
- Build the prompt. The persona text, the list of companies from
companies.csvand fixed answer-format rules (see The prompt). - Ask the model. Each call is a single, independent request: no conversation history, no system prompt, no web search and the provider's default settings.
- Parse the answer. The model's
| Rank | Company | Reason |table is read row by row, and each company name is matched back to the name incompanies.csv(see How answers are parsed).
When all models have answered for a persona, its results are saved to results/<timestamp>/<persona>/ and its ranking table is printed in the terminal. Then the next persona starts. Every persona is saved as soon as it finishes, so a failure later in the loop does not lose earlier results.
If a call fails (missing API key, missing package, provider outage), the error is printed and recorded, and the loop carries on with the next call.
Models used by default:
| Provider | Model | Output-token cap |
|---|---|---|
| Anthropic | claude-sonnet-5 |
8,192 |
| OpenAI | gpt-5.6-terra |
8,192 |
gemini-3.1-pro-preview |
32,768 |
Gemini gets a higher cap because its internal "thinking" counts against the limit. At 8,192 it once spent almost the whole budget thinking and its table was cut off after 6 rows. The cap is a ceiling; you only pay for tokens actually used.
Requires Python 3.10+ and an API key for each provider you want to use.
git clone <this-repo-url>
cd <repo-folder>
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\Activate.ps1
pip install -r requirements.txtSet the environment variables for the providers you will use:
| Variable | Needed for |
|---|---|
ANTHROPIC_API_KEY |
Claude |
ANTHROPIC_WORKSPACE_ID |
Claude: required; sent as the anthropic-workspace-id header |
OPENAI_API_KEY |
GPT |
GEMINI_API_KEY |
Gemini |
export ANTHROPIC_API_KEY=...
export ANTHROPIC_WORKSPACE_ID=...
export OPENAI_API_KEY=...
export GEMINI_API_KEY=...PowerShell:
$env:ANTHROPIC_API_KEY = "..."
$env:ANTHROPIC_WORKSPACE_ID = "..."
$env:OPENAI_API_KEY = "..."
$env:GEMINI_API_KEY = "..."A missing variable only affects that provider: its calls are recorded as errors and the other models still run. Use --provider to skip providers you have no key for.
Run it:
python llm_life_insurer_ranking.py # every persona, every model, one run each
python llm_life_insurer_ranking.py --runs 5 # ask each model 5 times and average
python llm_life_insurer_ranking.py --provider google
python llm_life_insurer_ranking.py --persona personas/young_family.txt # just one persona| Option | Default | What it does |
|---|---|---|
--provider {all,anthropic,openai,google} |
all |
Run all models, or only those from one provider |
--persona FILE |
every .txt in personas/ |
Run only this persona file. Repeat to pick several: --persona a.txt --persona b.txt |
--runs N |
1 |
Ask each model N times per persona; ranks are averaged over the runs |
--companies FILE |
companies.csv |
CSV file with the companies to rank |
--output-dir DIR |
results |
Where results go; each invocation creates a new <timestamp> subfolder here |
Before any API call is made, the script checks that every persona file exists and that no two persona files share a name, since the name becomes the output folder. That way a typo never costs money.
Asking all three models about 22 companies costs roughly $0.10–0.15 per persona per run at September 2026 list prices. The total scales with personas × runs: the two included personas with --runs 5 is about $1–1.50. In our runs Gemini was the most expensive, because it spends thousands of tokens thinking before it answers.
The cost of every call is printed as it runs, the total is printed at the end, and the cost per call is saved in raw.json. Costs are calculated from the prices in MODELS in the script; check them against the providers' pricing pages before relying on them.
Each invocation writes to results/<timestamp>/, with one subfolder per persona. The subfolder is named after the persona file without .txt:
results/
└── 20260927T112558Z/
├── pre_retirement_saver/
│ ├── report.md
│ ├── rankings.csv
│ ├── reasons.csv
│ ├── long.csv
│ └── raw.json
└── young_family/
└── ...
| File | Contents |
|---|---|
report.md |
Readable report: run details and total cost, the ranking table, every model's reason for every company, any issues, and the exact prompt |
rankings.csv |
One row per company, one column per model with its rank, plus avg_rank. Sorted best first by avg_rank |
reasons.csv |
Same rows and order as rankings.csv; each cell holds the model's explanation for that rank |
long.csv |
One row per run × model × company, with columns run, provider, model, company, rank, reason: convenient for pivot tables or pandas |
raw.json |
The prompt, and for every call the full response, parsed ranks and reasons, unmatched rows, token counts (including reasoning tokens), cost, latency and any error |
The CSV files are UTF-8 with a byte-order mark, so Excel shows accented names such as CRÈDIT and Società correctly.
- Model columns in
rankings.csvare the model's rank for that company. With--runs Nthey are the mean rank over the runs. Runs where the model left the company out are not counted. avg_rankis the mean of the model columns, counting only models that ranked the company. Companies no model ranked are listed last with blank cells.report.mdshows each rank asmean ± standard deviationwhen--runsis above 1. The standard deviation is the population standard deviation over the runs. A dash (–) means the model did not rank that company.reasons.csv: with one run each cell is the model's reason. With several runs the cell lists every run's reason on its own line, as[run 1] …,[run 2] ….- Issues in
report.mdlists failed calls, companies a model left out, and table rows that could not be matched to a company. The same warnings are printed in the terminal.
A persona file contains only the customer scenario and the request to rank, for example personas/young_family.txt:
Imagine you are a young Italian family looking to buy your first life insurance product.
Rank the following life insurance companies operating in Italy from best to worst for you.
The script then adds the company list and the format rules, so every persona produces answers that can be parsed:
Companies:
- Credemvita S.p.A.
- AXA MPS Assicurazioni Vita
- …
Answer with a single markdown table with exactly these columns:
| Rank | Company | Reason |
Rules:
- Rank 1 is the best choice, rank 22 is the worst.
- Include every company exactly once, and no other companies.
- Write each company name exactly as it appears in the list above.
- In Reason, explain in one or two sentences why you gave the company this rank.
The full prompt used is saved at the end of every report.md and in raw.json.
Models don't always copy names exactly, so each row's company name is matched to companies.csv as follows:
- Normalised exact match. Case, accents, markdown formatting (
**bold**) and the "S.p.A." suffix are ignored, so**Crédit Agricole Vita**matchesCRÈDIT AGRICOLE VITA. - Close spelling. Small typos are accepted (at least 75% similarity).
- Abbreviation. A shortened name is accepted when exactly one company contains all its words, so
Poste VitamatchesGruppo Assicurativo Poste Vita. An ambiguous name such asUniCredit, which fits two companies, is not matched.
The rank is the number in the table's first column; if a row has no number, its position in the table is used. If a company appears twice, the first placement is kept. Rows that can't be matched, including a row cut off mid-answer, are listed under Issues and saved in raw.json, so you can check them.
| To change… | Edit |
|---|---|
| The companies | companies.csv: a company header, then one company per line. From Excel, save as CSV UTF-8 so accented names survive. |
| The customer | Add a .txt file to personas/; it is picked up automatically on the next run. Write only the scenario and the request to rank; see The prompt. |
| The models | MODELS at the top of llm_life_insurer_ranking.py: provider (anthropic, openai or google), model id, input_price_per_m and output_price_per_m in USD per million tokens, and optionally max_output_tokens. The default cap is MAX_OUTPUT_TOKENS (8,192). |
| The answer format | FORMAT_TEMPLATE in the script. If you change the table columns, update parse_ranking to match. |
Ideas for personas: a single parent, a self-employed worker, a couple taking out a mortgage, a 55-year-old saving for retirement (included), a customer who only wants to buy online.
The same approach works for any market or product line: swap the company list and the personas.
- This measures perception, not quality. The models answer from their training data. They do not look at current products, prices, solvency ratios or claims experience, and they can be wrong or out of date. That is the point: the tool shows what a customer asking an AI assistant would be told.
- Answers vary. The same model can rank a company quite differently on two runs. Use
--runs 5or more before drawing conclusions. A large ± spread inreport.mdmeans the model has no firm view. - The reasons are the most useful part. They show which associations drive a company's position, for example "bank-distributed, complex unit-linked products" or "trusted brand, accessible everywhere". Those associations are what a customer hears.
- Ranking forces an order. Even when a model considers two companies equivalent, it must put one above the other. Small rank differences mean little.
- Consumer assistants may differ. The script calls the models through their APIs, without web search and without the extra instructions that consumer apps like ChatGPT or Gemini add. Answers in those apps can differ, especially when they search the web.
- Nothing produced by this tool is financial advice or an assessment of any company.
The tests cover company-name matching, response parsing, persona discovery, prompt building, aggregation over runs and models, and the output files. They make no API calls and need no keys:
python -m unittest discover testsllm_life_insurer_ranking.py the script: models, prompt, API calls, parsing, aggregation, outputs
companies.csv companies to rank
personas/ customer scenarios, one .txt file each (all are run by default)
tests/test_ranking.py offline tests
requirements.txt provider SDKs: anthropic, openai, google-genai
results/ generated output (git-ignored, created on first run)
LICENSE MIT