SM/RSoftware Marketing Resource Subscribe
Positioning/Analysis

Gorgias put its own AI agent in a public benchmark, and published every place it lost

A $100M-ARR customer-service AI vendor ran its own agent through the same public benchmark as 17 rivals and published the categories where it lost — a possible template for how B2B software marketers earn trust with AI-driven buyers.

Sienna McphersonSienna Mcpherson✓Contributing writer
Sep 26, 2026 · 4 min read
X in f
A black king chess piece standing upright in sharp focus while toppled white and black chess pieces lie scattered and blurred behind it
One piece still standing after the rest are down. Photo: George Becker / Pexels

Gorgias, a $100 million-ARR customer-service AI vendor, has published a running public benchmark that scores its own agent against 17 named rivals — and shows it losing outright in one of the two jobs the benchmark measures.

The report, at evals.gorgias.com, tests AI shopping and support agents on live Shopify storefronts rather than demos or vendor decks. More than 8,300 conversations across 212 stores are scored blind against a public rubric and refreshed weekly. Gorgias ranks first overall and ties for first in support. In pre-sale shopping, it comes second, and the report says so on its own page.

What changed

Gorgias tested itself under the same conditions as every competitor: each conversation opens a brand-new, unauthenticated session with nothing carried over from the last one, questions are typed out rather than picked from scripted prompts, and an auditor persona that never asks for a human — every handoff has to be triggered by the agent admitting it couldn't answer. Per SaaStr, whose founder Jason Lemkin wrote up the report on September 26, a separate reviewer checks each score afterward and has to back up the original judge's call roughly nine times out of ten, or the result doesn't stand.

The numbers aren't uniformly flattering. Gorgias's support agent resolves 74% of conversations without a human and scores 76 out of 100 on answer quality, second only to Sierra's 79 against a field median of 60. Its shopping assistant has the highest answer quality in the field, but SaaStr reports it lost the pre-sale composite to a competitor called Envive, 65 to 72, because shopping answers average 18.4 seconds versus Envive's 7.9 — with 28% of Gorgias's shopping answers running past 20 seconds. Gorgias weighted response speed at 25% of the shopping score, more than double the 10% it uses for support, a choice SaaStr notes cost Gorgias the top spot: scored under the gentler support weighting, it would have finished first in shopping too.

"a benchmark that hides its own is not worth reading"

Gorgias's own summary of the results, quoted in part above, doubles down on that logic. The benchmark also surfaces a pattern that has nothing to do with any single vendor: across the field, the most common failure isn't a wrong answer — it's an authentication wall, a repeated request for an order number, or a loop that ends in a human handoff. Configuration, not the underlying model, is what separates a vendor's best-performing store from its worst. And nearly a third of the "AI chat" widgets the benchmark's crawlers found on live stores produced no real conversation at all.

18vendors scored head-to-head
8,300+live conversations judged
212real storefronts, not demos

Who this affects

SaaStr's argument is aimed at any B2B software company selling into a market where buyers — increasingly, buyers' AI agents — can no longer be won with a features slide or an analyst quadrant. "Vendors that publish structured, checkable evals will get cited more," Lemkin wrote, arguing that a gated comparison PDF gives an AI agent nothing to read, while an open, versioned rubric gives it something to cite. That argument has a wrinkle worth disclosing: SaaStrFund led Gorgias's seed round, and SaaStr's own writeup says it "pushed" Gorgias toward publishing the evaluation — this is a portfolio company's homework being graded by its own investor, not a neutral trade-press review.

It's also worth noting the report currently blocks automated crawlers in its robots.txt, even though the underlying test harness is open-sourced on GitHub — a gap between the citation argument and the current setup that SaaStr itself flags as something Gorgias still needs to fix.

What a software marketer should do differently

The template SaaStr extracts from Gorgias isn't really about AI agents specifically — it's a positioning move any vendor with a testable product could copy, and it works whether or not an LLM ever reads the results.

  • Test the live product against named competitors, not a sandboxed demo against an anonymized "Competitor A."
  • Publish the rubric and the weighting before publishing the score, and say plainly that you wrote both.
  • Include the categories where a competitor wins — a single acknowledged loss is what makes every win next to it believable.
  • Refresh the results on a schedule and show the trend line, not a one-time snapshot that's easy to cherry-pick.
  • Open the methodology to the same crawlers — human and automated — that you want citing the results later.
What to do

Before commissioning another category-leader graphic, ask whether it would survive being rerun by someone who doesn't work for you. If it wouldn't, a buyer's AI agent will treat it the same way a skeptical human does: as marketing, not evidence.

None of this requires an AI product to attempt. A pricing page, an onboarding flow, or a support SLA can be benchmarked against named competitors on the same terms — tested live, scored on a public rubric, refreshed on a schedule. The mechanism Gorgias used is old: show your work, include the losses, let anyone check it. What's new is the audience checking it first is increasingly a model deciding who gets shortlisted, not a person reading a datasheet.

PositioningAI AgentsCompetitive IntelligenceB2B Marketing
Share: X · LinkedIn · Facebook
↓Read next

One briefing, every Tuesday.

The week in software marketing: the news that matters, one unsponsored review, and the numbers behind both.

Free · Unsubscribe anytime