Hey there

Welcome to my blog

How Well Calibrated Are Jev’s Probabilities?

Jev is an AI forecasting service that assigns probabilities to yes-or-no questions. Each question can include true and false text defining what counts as yes or no. With no deadline and both fields set to empty strings, Jev gave 35% to bitcoin will reach 100 million and 14% to bitcoin will reach 1 million, averaging 100 calls each. Calibration asks whether forecasts assigned 20% come true about one time in five across many independent events. ...

September 21, 2026 · 6 min · npow

The State of Robotics (September 2026)

A September 2026 census of 2,538 robotics and physical-AI companies across 24 use cases, with funding, hiring, geography, deployment, and an interactive company directory.

May 31, 2026 · 2 min · npow

Does Caveman Mode Actually Work?

Caveman is a Claude plugin that asks the model to answer in terse fragments. Does it save money? Sometimes—mostly when the model would otherwise produce a very long answer. I compared Caveman’s lite, full, ultra, and wenyan modes with two baselines: no added instruction and “Answer concisely.” The benchmark covered 15 prompts, three runs per condition, and two Claude Opus models. I measured token cost, not answer quality. Model and task Caveman cost vs. concise Claude Opus 4.7, short Lite +10%; Full +3%; Ultra +3%; Wenyan +10% Claude Opus 4.7, long Lite +9%; Full +10%; Ultra −5%; Wenyan −1% Claude Opus 4.6, short All modes: 3–16% less; uncertain Claude Opus 4.6, long Lite −54%; Full −56%; Ultra −59%; Wenyan −58% Positive numbers mean higher cost; negative numbers mean savings. The practical rule: Caveman has to save enough output tokens to pay for its added instructions. That is easy with a long tutorial and hard with a short answer. ...

April 20, 2026 · 4 min · npow

We Benchmarked MCP Against Code Generation. MCP Won (Mostly).

TL;DR: For a small, well-designed API (10 tools), structured MCP tool calls consistently outscore code generation on correctness — 0.99 vs 0.97 — and the gap concentrates in tasks where domain-specific logic matters. Adding a reference document to MCP tools costs 6–14% more tokens with zero accuracy gain. Cloudflare’s search+execute pattern matches MCP accuracy but uses more tokens. With MCP tools, Haiku is within 1% of Opus at 1/12 the cost. ...

March 18, 2026 · 12 min · npow

Finding the Human in the Machine

50+ open-source orgs are rebuilding how they evaluate contributors. Here’s what’s emerging.

March 6, 2026 · 8 min · npow