Bug0’s cover photo
Bug0

Bug0

Software Development

San Francisco, California 5,479 followers

About us

Bug0 | Meet your AI QA Engineer Bug0 delivers AI-native browser testing that runs itself. A complete managed QA service that eliminates browser testing struggles for web apps. AI creates the tests, our QA experts verify them, and bugs get caught automatically. We help teams achieve 100% critical flow coverage in 7 days with zero setup. Our system connects directly to your CI/CD pipeline, generates self-healing Playwright tests, and adapts automatically to UI changes. With the Forward-Deployed Engineer (FDE) model, Bug0 combines agentic AI with embedded QA experts who work like an extension of your product team. Each pod includes AI-powered test creation, human-in-loop validation, and managed infrastructure that runs 500+ tests in minutes. ✅ 100% critical flows in 7 days ✅ 80% total coverage in 4 weeks ✅ SOC 2 Type II ready ✅ Zero setup, connects to CI/CD directly ✅ Cancel anytime, keep all tests Bug0 helps modern engineering teams ship faster, catch more bugs before users do, and reduce QA overhead by up to 80%. Outcomes, not QA overhead. 🔗 bug0.com

Website
https://bug0.com
Industry
Software Development
Company size
2-10 employees
Headquarters
San Francisco, California
Type
Privately Held
Specialties
AI-Powered QA Automation, End-to-End Browser Testing, Automated Test Generation, CI/CD Integration, Web Application Testing, Regression Testing, Web App QA, Zero-Maintenance Testing, Developer Productivity, and Quality Assurance

Employees at Bug0

View 23 employees at Bug0

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

Updates

  • Bug0 reposted this

    AI software factory covers four different purchases, and vendors rarely tell you which one you're being shown. Buy a platform and you get the plumbing between the stages, then you integrate it into your own stack. Assemble one and you pick the sandbox, the coding agent, the verification layer and the credential broker separately. Self-host the lot if compliance says code cannot leave the network. Or have one built inside your environment and handed over to you. The stages are the same in every case: intake, isolation, implementation, verification, pull request. What changes is who does the integration work, and how long you wait for the first PR. Two things I'd check before sitting through a demo. Isolation is harder than it looks and gets underestimated in almost every plan. And assembled stacks tend to drop verification, which leaves you with an agent that opens pull requests nobody trusts. There is also no published first-attempt merge rate from anyone, so you cannot compare reliability across vendors yet. The tools available at each layer: https://lnkd.in/eVxMBMPY

    • No alternative text description for this image
  • Bug0 reposted this

    In 10+ years of building and testing software, the most expensive habit I've seen isn't bad code. It's the rerun. A CI job fails. Someone glances at it, decides it's "probably flaky," hits rerun. It passes. The PR merges. Fifteen seconds, no harm done. Except: run that for a 10-engineer team. 10 flaky reruns per engineer a day, 5 minutes of attention each. That's over 8 hours of collective attention every day, more than $160,000 a year, spent on tests nobody trusts. Slack's engineering team measured their flaky failure rate at 56.76% before they fixed it. Most teams have never measured theirs once. I wrote up the full math, why retries and quarantine lists fix nothing, and what removes the cause instead of hiding the symptom. Link in the first comment.

    • No alternative text description for this image
  • Bug0 reposted this

    When you pick an AI testing vendor, you also pick their AI model provider. Nobody puts that on the pricing page. Most AI testing tools call one provider's SDK directly. Which works fine until a model gets deprecated with a short migration window, or the provider has an outage during your deploy freeze, or a model update changes how your assertions get graded and your pass rate moves for no reason you changed. For scale: Claude 3.5 Haiku went from deprecation notice to scheduled shutdown in about six months. If your test suite was hardwired to it, that migration is now your project. At Bug0 we route every AI call in Passmark, our open-source engine, through an AI gateway instead (Vercel AI Gateway, OpenRouter, OpenCode Zen, Cloudflare). Provider trouble becomes a routing decision, not an incident. I wrote up how it works, plus the questions to ask any AI testing vendor before you commit. Link in the first comment.

    • No alternative text description for this image
  • Bug0 reposted this

    We verified 120 AI-built flows last month, and 70% of the bugs that made it into PRs came from the same three places. 1️⃣ About 45% were edge cases the prompt never mentioned. 2️⃣ Another 30% broke against existing code the AI never bothered to read. 3️⃣ The last 25% were assumptions the AI just made up on its own. That last group is the one that scares me, because you don't see it coming. The AI writes a function, it compiles, the unit test it wrote alongside passes too, and static analysis comes back clean. Looks done… Then you find out it called an API field that doesn't even exist on your production model. Nothing throws, nothing breaks, the feature just quietly does nothing. And you don't notice until a customer does. Unit tests were never going to catch this. They're checking the AI's logic against the AI's own assumptions, so of course they agree with each other. You only catch it when you run the real flow, with real data, against the real API. End to end. That's the gap our FDEs run into pretty much every week. They build the end-to-end coverage on Passmark, our AI engine, and it reruns and heals itself on every deploy so this stuff surfaces before the PR does, not after. We open-sourced Passmark too, so you can go poke at it. We wrote up the full QA testing playbook from what they actually see out there. Check below.

    • How to test AI-generated code before it ships.
  • Bug0 reposted this

    Every model since GPT-4 could write hundreds of unit tests effectively. That's not what's new with Claude Fable. What's new is that nobody's reading the code anymore. Fable ships the whole feature: implementation, tests, the lot. Earlier models wrote code you reviewed line by line. Now the diff is 4,000 lines, it arrived in twenty minutes, and reviewing it properly would take longer than writing it yourself. So teams skim it, see green tests, and merge. And before someone says "use a second model to write the tests": that fixes the wrong problem. Unit tests, whoever writes them, check that the code does what the code says. The bugs I've watched cost teams real money this year lived somewhere else entirely. Real auth meeting a real browser. A third-party API timing out at 2am while the checkout flow just... waited. None of that was in the diff. None of it ever is. The human checkpoint didn't get automated out of the pipeline. It got skipped. Someone still has to answer whether code nobody on your team read is safe to put in front of customers. We're seeing this play out in our pipeline at Bug0. The fastest-growing demand isn't for another testing tool, it's for our forward-deployed engineers: teams shipping AI-written code who want an actual human, outside the codebase, accountable for "safe to ship." The faster a team has gone with AI, the sooner they call. Full essay on where the QA work went: https://lnkd.in/grq6BTfn

  • Bug0 reposted this

    Anthropic says 90% of its code is now written by AI. Microsoft and Google are at 25–30% and climbing. The part that didn't make the headline: the bugs are climbing too. 📈 The 2026 reports are blunt. 43% of AI-generated changes still need debugging in prod, after they passed QA and staging. Incidents per merged PR are running about 3x the pre-AI baseline. We're writing code faster than ever. We're shipping bugs faster too. And the one role meant to catch them is the one nobody owns. Open any mature codebase and you'll find a test that's been skipped for months. A TODO sits above it. Whoever wrote the TODO doesn't work here anymore. QA is everyone's responsibility, which means it's no one's job. So the suite rots. One skipped check, one "we'll fix the flake later," and the team trains itself to scroll past red. Then something breaks in prod that a test you already wrote would have caught. It was just turned off. This isn't a tooling problem. It's an ownership problem. And AI just made it 3x more expensive. Most teams can't justify a full-time QA hire. So we built the owner instead. Bug0's engineers run your coverage end to end: planning, writing, triage, keeping it green. Passmark, our open-source engine, does the heavy lifting underneath. https://lnkd.in/gpUu5gm7 How it actually works: 1. Steps are written in plain English and compiled to Playwright. Every passing step is cached, so runs stay fast, cheap, and deterministic. When the UI changes and a step breaks, the model re-resolves it against intent and re-caches it. That's the healing. No brittle selectors to maintain. 2. Assertions never trust one model. Claude and Gemini each judge them, and a third breaks ties. We check against the video of the run, not just the final DOM, so a success toast that flashes for half a second still counts. 3. Then a human reviews every run and decides what's a real bug. The engine gives leverage. The engineer gives judgment. That second half is the 43% the tools keep missing. Passmark is closing in on 1,000 stars 🌟, the fastest-growing AI Playwright library on GitHub. 100% coverage in four weeks. 0% flake. Less than the cost of one QA hire. The AI writes the code now. Someone still has to own whether it works.

    • Line-and-bar chart, 2023–2026: AI-generated code (line) rises ~5→90, outpacing production incidents (bars) ~20→60.
  • Bug0 reposted this

    Running a model on every browser action: 4 to 8 seconds of latency per step, $0.02 to $0.05 in tokens. At 50 steps across 200 tests per nightly run, that's $200 to $500 a night just in inference. Across environments and branches, the bill exceeds the QA hire it was supposed to replace. Every AI testing tool faces the same paradox. The LLM is what makes it work. The LLM is also what makes it economically impossible. A cheaper model doesn't fix this. A different architecture does. One where the LLM isn't on the hot path. We've been building toward this for 9 months. The Bug0 core is now open source as Passmark (https://lnkd.in/gpUu5gm7, 800+ ⭐). Three primitives: 1. Cache-first execution. The first successful run resolves an action and caches it in Redis. Every subsequent run replays from cache. Zero AI calls until something breaks. 2. Auto-heal on cache miss. When a cached action fails because the DOM drifted or the UI got refactored, the AI re-resolves and updates the cache. The test self-repairs. 3. Multi-model consensus on assertions. A single model hallucinates. So we run Claude and Gemini in parallel. If they agree, the assertion ships. If they disagree, an arbiter model decides. Hallucination becomes a probability you design around. You don't wait for a vendor to solve it for you. There's a fourth piece. For the small fraction of steps that need pixel-level reasoning (sliders, canvases, drag-and-drop) you can flip to OpenAI's computer-use agent for that one step. Cheap snapshot path and expensive visual path inside a single test. Most tools force you to pick one upfront. The principle: the LLM is a coprocessor, not a CPU. Every browser testing tool will converge on this. The ones that don't won't survive a budget review next year. Bug0 is a forward-deployed engineering team running on top of Passmark. We embed senior engineers with you, bring the infra (the executor, the caching layer, the model orchestration), and own the outcome. 100% coverage of your critical user flows in weeks, not months. The suite, the maintenance, the on-call: all on us. If you're the engineering leader who'll own this decision in the next 12 months and want to walk through the architecture, not the sales pitch, that's bug0.com/join-pilot.

    • GitHub repo screenshot of Bug0's open-source Playwright AI library called Passmark. Optimised for AI powered QA end to end testing.
  • Bug0 reposted this

    Shipped passmark 1.0.14 today. We kept hitting the same problem: screenshots are great for static UI, but they completely miss the stuff that flashes and disappears. Toast notifications, snackbars, transient error states - the agent would take a screenshot a half-second too late and the assertion would silently pass (or fail) for the wrong reason. So, we added video assertions. Set `video: true` on any assertion in runSteps, and passmark records the full test execution, and evaluates the assertion against the entire recording. The model sees everything that happened, not just one frame. A few details:  - Recording spans the whole runSteps call  - Local file is deleted after the model consumes it  - Default dir is /tmp/passmark-recordings, override via configure({ videoDir })  - As of now, we call gemini under the hood. So, you need to set `GOOGLE_GENERATIVE_AI_API_KEY` if you want to use video assertions. This was unblocked by page.screencast API landing in Playwright 1.59. Big thanks to the playwright team. If you're writing e2e tests using Passmark and always wondered if your assertions are silently missing the toast that flashes for 1s, then this update is for you. https://passmark.dev

Similar pages