Expected Parrot (YC F25)’s cover photo
Expected Parrot (YC F25)

Expected Parrot (YC F25)

Technology, Information and Internet

Cambridge, MA 1,279 followers

Design, run, and analyze AI persona studies, then validate results with human respondents—all in one workflow.

About us

Design and launch surveys and automated interviews with AI and humans, all in one place.

Website
https://www.expectedparrot.com
Industry
Technology, Information and Internet
Company size
2-10 employees
Headquarters
Cambridge, MA
Type
Privately Held

Employees at Expected Parrot (YC F25)

View 3 employees at Expected Parrot (YC F25)

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

Updates

  • Expected Parrot (YC F25) reposted this

    There's a truism in economic research that one should "trust literatures, not papers" but doing good literature reviews is *very* time consuming. This is especially the case if you are trying to stay on-top of a fast moving area like the effects of AI on the labor market. Alex Imas and Jacob Schaal has a great blog post (link in the comment) summarizing and analyzing this literature. This morning, I had our Expected Parrot (YC F25) research agent---which is built to assist with all parts of the research process, including literature reviews---build on their work by collecting all the papers cited and applying the findings to a new report on NYC AI impacts. I then had Fable 5.1 make a video summarizing both the literature review and the NYC report.

  • Expected Parrot (YC F25) reposted this

    Pew just released a study in which their "digital twins" <https://lnkd.in/gZtpcEmW> (Claude Opus 4.6 personas built from each panelist's demographics and political typology answers) missed those same people's real answers by about 12 points on average. But if the goal is the toplines, you don't need to model individuals at all. I gave the same model, Opus 4.6, all 288 questions from Pew's three waves and asked it to predict each topline distribution directly, in one shot. Error fell by half, from 12.5 points to 6.1. That comparison is clean on contamination. Opus 4.6's training data ends in August 2025, months before these surveys were fielded, so it couldn't have seen the answers. Newer frontier models did even better, down to 4.2 points for Claude Fable 5.1, which beat the twins on 93% of questions. Their training data may include Pew's published results, though, so treat those numbers as an upper bound on what's possible. I ran the whole thing with Expected Parrot (YC F25)'s research agent (disclosure: I'm a co-founder), from extracting the questions out of Pew's PDF to the final report. It took about 45 minutes and cost about $19 in model calls. Link to study in the comment. Full report, code, and data in comment.

    • No alternative text description for this image
  • Expected Parrot (YC F25) reposted this

    Traditional market research is in flux, with problems such as low response rates, long wait times for answers, and low quality data. AI can help. MIT Sloan associate professor John Horton makes the case that large language models can effectively deliver informal market insights that can later be refined through further investigation and more formal research practices. “Companies might build and test three or four variants of something, but they’re not going to test 700,000,” Horton said during a recent talk. “Being able to preprocess and surface problems is really powerful.” Read Horton’s recommendations on how to help organizations maximize the value of AI-enabled market intelligence without getting led astray: https://lnkd.in/ewhiv_ZR

    • Photo of MIT Sloan associate professor John Horton with quote: "Companies might build and test three or four variants of something, but they're not going to test 700,000. Being able to pre-process and surface problems is really powerful."
  • Expected Parrot (YC F25) reposted this

    Pew just released a study in which their "digital twins" <https://lnkd.in/gZtpcEmW> (Claude Opus 4.6 personas built from each panelist's demographics and political typology answers) missed those same people's real answers by about 12 points on average. But if the goal is the toplines, you don't need to model individuals at all. I gave the same model, Opus 4.6, all 288 questions from Pew's three waves and asked it to predict each topline distribution directly, in one shot. Error fell by half, from 12.5 points to 6.1. That comparison is clean on contamination. Opus 4.6's training data ends in August 2025, months before these surveys were fielded, so it couldn't have seen the answers. Newer frontier models did even better, down to 4.2 points for Claude Fable 5.1, which beat the twins on 93% of questions. Their training data may include Pew's published results, though, so treat those numbers as an upper bound on what's possible. I ran the whole thing with Expected Parrot (YC F25)'s research agent (disclosure: I'm a co-founder), from extracting the questions out of Pew's PDF to the final report. It took about 45 minutes and cost about $19 in model calls. Link to study in the comment. Full report, code, and data in comment.

    • No alternative text description for this image
  • Expected Parrot (YC F25) reposted this

    A good example of why most marketing teams are missing much of the value agents can create. John Horton and the team at Expected Parrot (YC F25) have made research surveys version-controlled. Which means every question, logic change and alternative version has a visible history. You can compare what changed, test different approaches without overwriting the original, catch errors and reproduce the research later. I love nerdy stuff like this when its actually useful. Coding agents handle the technical plumbing, so marketers get the benefits without learning Git. Now, why should a founder, head of marketing or product marketer care? Because an agent is only as useful as the system you give it to work with. Give it a page and ask, “How can we improve this?” and you’ll get slop opinions. Give it structured buyer evidence, a controlled research method, versioned inputs and clear evaluation criteria, and it can help you investigate what your ICPs actually understand, trust and care about. That’s also where synthetic research becomes really useful: it doesn’t mean inventing a few generic AI personas and asking whether they “like” your homepage. It means creating simulated buyers grounded in real ICP and voice-of-customer research, then using them to test specific questions before a human makes the final call. In our teardown and messaging work, we use that principle + EP to: → Ground synthetic respondents in real customer interviews, sales calls, objections and buying triggers. → Define what the page or message needs the right buyer to understand. → Test what those buyers notice, remember, trust, question and would do next. → Compare messaging directions without losing the original research or rationale. → Turn repeated patterns into prioritised fixes, then review them against the source evidence before recommending anything. I would never replace marketers with agents, but agents can help if we give them a research system that helps us make better decisions for the people we're trying to reach.

  • Expected Parrot makes it easy to get useful feedback from synthetic personas before you launch, helping teams spot confusion, pressure-test assumptions, and identify what to explore further. Love this example from Chris Silvestri!

    I tested a specialist B2B consultancy homepage with 18 synthetic buyers. Here's what I've found... The consultant helps established companies win complex, high-value work. The homepage needed to signal specialist authority, communicate a clearer set of offers, and stay sober enough for commercially sceptical buyers who hate marketing speak. I was super curious to see what these AI personas “thought” about it. All 18 classified the business as a specialist rather than a generalist. All 18 read the voice as sober authority, not dry copy. And the proof he used made the positioning feel credible too. So far so good. Surprisingly, what respondents didn’t really get was the mechanism, the what he does and how he does it. They understood who the consultant helps and the commercial outcome, but they could not yet picture what the system actually looked like. It’s a typical case of not being clear on your jobs to be done. The consultant was so focused on making sure he didn’t sound scammy, that he forgot to clarify how he actually helps his clients. But this is why synthetic validation pre-launch should be mandatory now that the tools are out there. The panel we created (with Expected Parrot (YC F25)) showed which interpretation repeated across buying contexts, which objections came up together, and even used two different models to spot any differences. Synthetic buyers are not customer truth, but , especially in B2B, they’re super handy to test comprehension, expose objections, and decide what we should dig deeper into with more customer interviews or live experiments. Speaking of, next week I’m running an async homepage teardown with Exit Five using the my own process and adding on top this synthetic validation round. (link in the comments if you want to join the community and submit yours)

    • No alternative text description for this image
  • Expected Parrot (YC F25) reposted this

    How do you define product-markets that don't exist yet? John Horton and I built a method for it using LLMs Expected Parrot (YC F25). Check it out!

Similar pages

Browse jobs