Study protocol · version 1
How we'll test the chat on 1,265 B2B websites
This is the plan for the Speed-to-Lead Index, fixed before data collection began. We're publishing it first so you can hold us to it: what we'll measure, how, on whom, and how we'll count the results. Before the plan was final we did test our tools on 92 websites, and 42 of those companies are in the sample. Those visits shaped the tools, not the findings. Every company is visited again under this plan, and no test data goes into the results. If anything changes after data collection starts, you'll find it at the bottom of this page, dated, with our reason.
- Frozen
- 22 September 2026, before any data collection
- Amended
- 23 September 2026, still before data collection (A1–A5)
- Checksum
SHA-256 16ae67a8000ca62cc2a2ca266c075288edbff8d54079d086030a5e8337ca85d0of the full internal protocol, which will be published with the report
Why this study
Ask anyone in sales how fast you should answer a new lead and you will probably hear a number that traces back to one study: The Short Life of Online Sales Leads, published in Harvard Business Review in 2011. It is still quoted everywhere, fifteen years and one AI wave later.
Source: The Short Life of Online Sales Leads, Harvard Business Review, 2011
A lot has changed since. Most B2B websites now have a chat box, and many of those boxes are answered by software rather than people. Nobody has gone back and checked what a buyer actually gets today when they land on a company's site with a question. That is the gap this study fills.
It is a descriptive study. We report what a visitor experiences, site by site and in aggregate. We do not claim that chat causes revenue, or that one chat vendor causes better results than another.
Predictions
Six predictions are fixed in advance. Whatever the data shows, we will report each of them as confirmed or not. Everything else in the report will be labelled exploratory, so you can tell the difference between what we set out to test and what we noticed along the way.
- H1 AI chats reply faster than chats answered by people (median time to first reply).
- H2 AI chats answer the pricing question less often than people do.
- H3 S&P 500 companies let a visitor type a question less often than Y Combinator startups.
- H4 Fewer than one in five sites with chat greet you on the pricing page with a line about pricing.
- H5 Fewer sites show chat after hours than during business hours, mostly because staffed chats switch off.
- H6 Fewer sites show chat on a phone than on a desktop.
Who we study
You'll find about 1,265 companies in the study, drawn from ten public lists. The S&P 500 shows how the biggest listed firms behave; YC and Techstars show what this year's startups do; the rest sit somewhere in between. Nobody is counted twice. Once a company has been drawn for one list, the later lists skip it.
| List | Companies | Source |
|---|---|---|
| S&P 500, companies selling mainly to businesses | 100 | Current index constituents |
| Forbes Cloud 100 | 100 | 2025 list |
| G2 Best Software, sales and marketing | 75 | G2 2026 Top 50 Sales and Top 50 Marketing |
| Deloitte Technology Fast 500, North America | 100 | 2025 list |
| Forbes 30 Under 30 companies | 150 | US and Europe 2026 lists |
| Y Combinator | 150 | The last eight batches |
| Techstars | 100 | Programs since 2019 |
| Venture portfolios | 240 | 30 each from a16z, Sequoia, Accel, Index, Lightspeed, Bessemer, Point Nine and Redpoint |
| B2B service firms | 150 | SelectedFirms US directories |
| Agency partners | 100 | HubSpot Elite and Diamond partners; Webflow Certified and Enterprise partners |
| Total | 1,265 | Random seed 20260923 |
We saved every list exactly as we found it, along with where it came from and the day we collected it, so you can check it. The draw inside each list follows a random order fixed by a published seed: rerun it and you get the same companies. Some sites turn out to be dead, consumer-facing or only in another language. When that happens, the next company in that list's order takes the place, and the swap goes in the log.
How a visit works
Every page is loaded in a fresh browser with no cookies, from a single location in the United States, in English. We then behave like an unhurried visitor. Wait five seconds, scroll halfway, wait again, scroll to the bottom, and stay for thirty seconds in total.
Cookie banners are left alone unless one covers the corner where chat launchers sit. In that case we pick the most private option on offer, usually reject or necessary-only, and carry on. If a site shows a bot check instead of its page, we try once more after a pause and otherwise record it as blocked. A blocked site never counts as a site without chat.
Stage one: looking, without typing
The first stage is read-only. On up to five pages per site, we look at whether chat appears, what kind it is, and what it says first, without typing a word. We open the widget the way a visitor would, by clicking its launcher, and we may press one button that leads to asking a question. That is as far as it goes.
The five pages are the homepage, the pricing page, one product page, one blog article and the contact or demo page, each found through the site's own navigation. The homepage is loaded twice. If it greets us differently each time, the site is rotating or testing its opening line, and we do not credit it with tailoring the message to the page.
Each site is checked on desktop during business hours in its own time zone, then again after hours and on a phone. That lets us see whether chat disappears at night or on mobile, which is where a lot of buyers actually browse.
The conversation
The second stage is the only one that sends messages, and only to sites where a visitor can type a question. Every company gets the same four questions, in the same order, typed at the same speed:
- “What does pricing look like for a company of about 50 people?”
- “Do you work with companies like ours, a 50-person B2B software company?”
- “Am I chatting with a person or an AI?”
- “Can I book a call with your sales team this week?”
How long do we wait? Up to ten minutes for the first answer, since a real person may need a moment to pick up the chat. After that it's three minutes per question. Once a reply has stopped changing for six seconds, we wait five more and send the next one.
A conversation ends early in two situations. If the chat asks for an email, phone number, name or company, we stop and record where it asked. If it shows real calendar slots, we stop and record how long it took to get there. We never pick a slot.
What we never do
We never type an email address, phone number, name or company into any chat, never submit a form and never book a meeting. Fake bookings would waste a real salesperson's time, which is the opposite of what this study is about. Time to meeting therefore ends at the moment a bookable slot appears.
If a person answers, our last message tells them the truth: that the conversation was part of a study on how B2B websites respond to buyers. Nothing personal about anyone who answers is published.
What we report
- Whether chat appears, and which product runs it
- What kind of chat it is: AI agent, staffed chat, buttons only, customers only, or asks for contact details first
- Whether a visitor can type a question
- The first thing the chat says, and whether it fits the page
- Time to the first reply
- Who replied: AI, a person, an automatic message, or nobody
- Whether the pricing and fit questions got real answers
- Whether an AI says it is an AI when asked
- Time until a bookable slot appears, and how booking works
Checking our instruments
A study like this is only as good as its instruments, so we check them before relying on them and report how well they did.
Two people label sixty test websites by hand, independently: whether chat is visible, what kind it is, and whether a visitor can type a question. We report how often they agree, settle the disagreements, and then require our scanner to match the settled labels on at least 95% of sites before it touches the study list.
To measure our own timing error, we built a test page whose chat replies after known delays and ran the conversation tool against it. On twelve replies between 1.5 and 60 seconds the measured time matched the page's own record to the millisecond.
Some calls need judgement, like whether an opening line refers to the page you're reading. Those go to a language model with a written rubric in front of it, and with the company's name and web address blanked out, so a famous brand can't sway it. We then pull one label in ten and check it ourselves. And a week after the main pass, we go back to a random tenth of the sites to see how much changed.
How we count
Every percentage comes with a 95% confidence interval. Reply times are reported as medians, plus the share of chats that answered within ten, sixty and 120 seconds. A chat that never replied is counted as still waiting when we stopped watching, not as having replied at the cutoff, which is the standard way to handle a clock that runs out.
We only report a separate figure for a group, such as a vertical or a fund's portfolio, when it holds at least forty companies. The six predictions are tested with a correction for testing several things at once. Differences between chat vendors are adjusted for company size, list and vertical, because larger companies tend to buy different tools, and they are described as associations rather than effects.
Our interest in the result
We should be upfront: Darwin sells AI chat to B2B websites, so we're not neutral about how this turns out. Hence the safeguards. The plan is public before any data is collected. The model that labels chats never sees a company name. You'll get the full dataset and the code with the report. And neither Darwin nor any of its customers appears in any ranking.
A week before publication, every company we name will receive its result and a way to tell us if we got something wrong. Corrections will be logged and published with the report.
Limits
The lists are named lists, not a random sample of every B2B company in the world. Everything is seen from one location in the United States, and some tools greet visitors differently depending on the company they appear to work for, so we measure what an anonymous visitor sees. Each site gets one conversation per condition, and the one-week revisit shows how much a single visit can vary.
The same companies will be checked again next year. Companies that close, get acquired or stop selling to businesses are replaced from the same list and flagged, so year-on-year changes are measured on the companies present both times.
Changes before data collection
All made on 23 September 2026, before data collection began.
- A1 · Lists
- A few lists changed before we started. You won't find the Inc. 5000 here: inc.com blocks automated collection, so Forbes 30 Under 30 companies took its place. Redpoint stands in for First Round, whose site asks AI agents to stay away, and SelectedFirms for Clutch, which sits behind a bot check like most directories of its kind. Two lists are a year old, the Cloud 100 and the Deloitte Fast 500, because the 2026 editions weren't out yet. And we read G2's lists from the Internet Archive, since g2.com shows a CAPTCHA.
- A2 · Visits
- Rules refined on test sites: a cookie banner over the chat corner gets its most private option; a teaser bubble counts as a launcher; one question button may be pressed.
- A3 · Test set
- How the test labels were made, and that sites which block automated browsers are left out of the test set.
- A4 · Sampling
- Forbes 30 Under 30 includes the Europe list; each company is drawn once; G2 yields 75 companies; phones use the same rules as desktops.
- A5 · Cookie consent
- Hand-labelling turned up something we hadn't planned for: some sites only load their chat after you accept cookies. Our registered visit declines, so those sites count as having no chat, and that stays the headline measure. On top of it we're adding a second, separate pass that accepts the cookie banner. That lets us tell you how many sites hide their chat until you say yes. We also taught the scanner to recognise Pipedrive's own chatbot.