SAN FRANCISCO — In the hyper-competitive landscape of Business-to-Business (B2B) artificial intelligence, every software vendor claims to offer the most accurate, efficient, and capable agent on the market. Historically, however, very few have been willing to show the underlying data. While the developer community has grown accustomed to rigorous, standardized evaluations (“evals”) for foundational Large Language Models (LLMs), the application layer—particularly customer experience (CX) and autonomous agents—has long relied on polished marketing decks, static feature checklists, and once-a-year analyst reports.
That paradigm is facing a radical disruption.
Gorgias, a $100M Annual Recurring Revenue (ARR) leader in e-commerce customer support and a portfolio company of SaaStrFund, has upended traditional B2B marketing playbooks. By open-sourcing its entire evaluation harness and publishing a comprehensive benchmark pitting its own agent against 17 competing vendors across more than 8,000 live conversations, Gorgias has established a new gold standard for software transparency.
The move highlights an urgent reality for the software industry: in an era where AI agents increasingly build vendor shortlists and inform executive decisions, buyers no longer tolerate generalized marketing claims. They demand verifiable, granular, and checkable data.
Main Facts: Deconstructing the Gorgias Benchmark
The core of the announcement centers on a newly released, public AI agent benchmark published by Gorgias. Unlike traditional vendor comparisons that rely on self-reported features or simulated environments, this initiative tests live products against named competitors in real-world scenarios.
Key Metrics of the Benchmark:
- Scale: Evaluates 8,356 live conversations.
- Scope: Tests 18 distinct software vendors in the e-commerce customer support ecosystem.
- Methodology: Completely open-sourced evaluation harness, allowing external auditing of rubrics, weights, and scoring logic.
- Transparency: Includes categories where Gorgias ranked behind competitors, validating the authenticity and objectivity of the testing framework.
While Gorgias secured the number-one spot in overall support capabilities within the benchmark, the company transparently highlighted competitor strengths—such as Yuma’s leading automation metrics and Envive’s response latency. By acknowledging these shortcomings rather than burying them, Gorgias fundamentally altered how enterprise software evaluations are perceived by prospective buyers.
Chronology: From Private Doubt to Public Accountability
The shift toward open-source evaluations did not happen overnight. It represents the culmination of a broader industry evolution regarding trust, validation, and the unique volatility of generative artificial intelligence.
- Early 2020s (The Feature Checklist Era): B2B software purchasing relied heavily on static assets. Buyers reviewed G2 grids, Gartner Magic Quadrants, and vendor-supplied PDFs claiming specific performance multipliers.
- The Rise of LLM Evals (2023–2025): As foundational models proliferated, the AI community demanded rigorous benchmarks (such as MMLU or HumanEval) to measure raw model intelligence. However, application-layer software agents remained insulated from similar public scrutiny.
- Late 2025 (The Catalyst for Change): Recognizing a growing skepticism among tech-forward e-commerce leaders, investors and operators began pushing for radical transparency. SaaStrFund, which led Gorgias’s seed round, actively encouraged the company to build and publish the most honest, detailed evaluation in the CX software space.
- September 2026 (The Public Release): Gorgias officially launches its AI Agent Benchmark report alongside an open-source GitHub repository (
gorgias/ai-agent-benchmark), sparking widespread industry debate on X (formerly Twitter) and within enterprise tech circles led by figures like SaaStr founder Jason Lemkin.
Supporting Data: Why Traditional Testing Fails AI
To understand why Gorgias’s approach is groundbreaking, industry analysts point to the fundamental difference between traditional SaaS architecture and generative AI agents.
Traditional Software vs. Generative AI Agents
- Determinism: A traditional CRM performs the exact same programmatic action every time a user clicks a button.
- Stochastic Behavior: An AI agent operates probabilistically. It might deliver an exceptional, nuanced customer service resolution on Monday, yet hallucinate or fail a routine compliance check on Thursday. Furthermore, silent backend model updates by foundational providers can degrade an agent’s performance overnight without warning.
Because of this inherent volatility, legacy evaluation tools are obsolete:
- Analyst Quadrants and G2 Grids: Updated infrequently and rely primarily on subjective user reviews rather than continuous, automated execution testing.
- Feature Checklists: Merely indicate whether a capability exists, providing zero insight into how well or how reliably the agent executes that capability under pressure.
- Vendor Demos: Heavily curated “happy path” scenarios designed to showcase strengths while masking edge-case failures.
By contrast, an active eval subjects the AI to real queries, captures every output, and grades performance against a versioned, written rubric. Gorgias’s implementation proves that continuous testing of real-world interactions is the only metric that truly matters to modern buyers.
Official Responses and Industry Reactions
The release of the benchmark has sent ripples through the B2B SaaS and AI communities, drawing praise from investors, founders, and enterprise buyers alike.
In a widely shared commentary, SaaStr founder Jason Lemkin underscored the importance of embracing vulnerability in competitive benchmarking:
"Everyone should publish the deepest, most direct competitive evals they can. Will there always be some bias? Yes. But Gorgias did it the right way in ecomm CX: open-sourced the whole harness, tested thousands of live conversations across numerous vendors, and showed where competitors excel."
Industry observers note that while Gorgias inevitably wrote the rubric and selected the evaluation weights—introducing a degree of structural bias—the decision to open-source the underlying code base neutralizes criticism. Because the harness is publicly available on GitHub, any competitor or skeptical buyer can inspect the code, modify the weights, and run the tests independently.
Implications: The Future of B2B Software Procurement
The widespread adoption of transparent, open-source evaluations carries profound implications for how software is marketed, evaluated, and purchased in the coming decade.
1. The Death of Gated Marketing Assets
B2B buyers have grown adept at filtering out promotional noise. When a vendor publishes a report acknowledging that competitors outperform them in specific niches, it builds immediate credibility. Buyers are far more likely to trust a company’s claims about its strengths when those claims sit alongside verifiable acknowledgments of its weaknesses.
2. Streamlining Enterprise Due Diligence
No mid-market or enterprise e-commerce brand has the internal bandwidth to test 18 different AI agents across thousands of live scenarios. By shouldering the massive logistical burden of running these extensive evals, Gorgias has effectively performed the market’s due diligence. Consequently, their benchmark report transitions from a marketing brochure into an indispensable industry standard.
3. The Rise of AI-Driven Shortlisting
An emerging development in enterprise procurement is the delegation of vendor research to AI agents. When a chief technology officer asks Claude or ChatGPT to identify the top three e-commerce support agents, those models look for structured, verifiable data. Gated PDFs, vague landing page claims, and unverified vendor slides are routinely ignored. Vendors that publish clean, versioned, machine-readable evaluation rubrics and open-source repositories will increasingly dominate AI-generated shortlists.
4. Internal Accountability and Product Velocity
Beyond external marketing benefits, publishing public benchmarks alters internal engineering culture. When metrics regarding response latency (such as Envive’s 7.9-second benchmark) or resolution accuracy are public and updated weekly next to named competitors, internal complacency vanishes. Teams are compelled to iterate faster and maintain rigorous quality control.
Best Practices: How B2B Companies Should Adapt
For software leaders looking to emulate the Gorgias playbook, industry experts recommend a disciplined, multi-step framework:
- Commit to Radical Honesty: Include named competitors and publish categories where your product loses. Hiding weaknesses destroys credibility; highlighting them builds trust.
- Open-Source the Harness: Do not just publish the final score numbers; publish the underlying code, datasets, and testing frameworks so your methodology can be independently audited.
- Version Your Rubrics: Maintain a clear version history in public repositories (such as GitHub) to ensure that scoring criteria cannot be quietly altered to hide performance gaps.
- Optimize for Machine Readability: Ensure that evaluation reports and technical documentation are structured so that autonomous AI research agents can easily parse, verify, and cite them.
As autonomous agents continue to reshape software evaluation and procurement, the era of the opaque marketing deck is coming to a close. Companies that embrace transparent, verifiable benchmarking will capture the trust of both human buyers and the AI systems that guide them.
