I priced enrichment for an AI agent at scale. The data costs more than the inference.
Last week I priced inference for an autonomous agent at $27 per active user per month. This week I priced the prospect data the same agent needs. Data wins by a factor of two — and unlike inference, the price isn't dropping.

Last week I priced inference for a realistic autonomous AI agent at roughly $27 per active user per month. The math that explains why headline LLM rates and the actual cost of running an agent are an order of magnitude apart.
This week I priced the prospect data the same agent needs to find anyone to message. Data won by a factor of two, and unlike inference, the price isn't dropping.
This is constraint B of the autonomous-AI-runs-your-company audit: the data wall. It's the constraint that doesn't dissolve with cheaper models. It's the constraint that quietly determines which autonomous-agent businesses survive 2026. And it's the constraint nobody pitches you on, because the answer is uncomfortable for anyone selling "AI runs your whole company."
Why high-quality B2B data is gated
The autonomous-agent pitch implicitly promises end-to-end customer acquisition: the agent finds the prospects, qualifies them, drafts the outreach, sends it, handles replies. The first step (finding prospects) lives or dies on the quality of the data the agent can access.
Where does high-quality B2B prospect data live in 2026?
- LinkedIn: behind authentication, with active enforcement against automated access.
- Apollo, ZoomInfo, Lusha, Clay, Cognism, Lead411: behind paid API access, credit-metered.
- Crunchbase, PitchBook: behind subscriptions starting in the low thousands per year.
- G2 intent, 6sense, Bombora: behind enterprise contracts that aren't quoted on websites.
The pattern is universal because the economic logic is simple. Verified contact data costs money to collect (research, enrichment, validation, decay-monitoring), and the moment it's free, it loses commercial value almost immediately. Gating isn't a bug. It's the only structure that keeps the data scarce enough to charge for. Any vendor that opens the gate destroys their own business. Any "free" data layer that scales also commoditizes itself out of usefulness.
So when an autonomous agent promises to find prospects, the question becomes: paid data or scraped data?
What paid data actually costs at scale
Let's price out an autonomous agent that qualifies 1,000 prospects per active customer per month, a modest number if the agent is going to send 50-100 outbound messages and you want any kind of qualification funnel.
Apollo lists at roughly $0.05–$0.10 per email-only contact and $0.30–$0.80 per contact when you include a verified mobile number, via a credit system metered per contact touched. Overage credits run $0.20 each on a $50 minimum top-up.
At 1,000 prospects per active user per month, email-only:
- Low end: 1,000 × $0.05 = $50 per active customer per month
- High end: 1,000 × $0.10 = $100 per active customer per month
For a product with $57 ARPU, that's 88-175% of revenue, before inference, before ads, before infrastructure, before paying the founder. The headline price is "starting at $59/month" for an Apollo seat; the actual cost of using it the way an autonomous agent needs to use it is an order of magnitude more.
ZoomInfo is materially more expensive. Their prospecting API starts around $50,000 per year; enrichment-only API access starts around $5,000 per year via the HubSpot marketplace. Full enterprise deployments run $30,000–$60,000+ per year, with per-seat add-ons at $1,500–$2,500 per user per year. ZoomInfo doesn't publish per-contact rates because most contracts are negotiated against credit pools, but the implied per-contact cost is consistently several multiples of Apollo's.
So the data math, at autonomous-agent scale, is roughly:
- Apollo, conservatively: $50–100/user/month
- ZoomInfo or premium providers: easily $100–300/user/month at full enrichment
- Specialty intent data layers (Bombora, 6sense): typically priced as annual contracts that don't decompose cleanly to per-user math, but materially add to the bill
Combine that with the $27/user/month inference cost from last week's post, and the autonomous agent's compute + data bill is $77–127 per active customer per month before any other cost. At a $57 ARPU, you're underwater on every customer, every month.
This is why nobody actually does this. The autonomous-agent platforms quietly skip the paid-enrichment path and reach for the scraping workaround instead.
The scraping alternative and its quality cliff
Scraping is cheaper. You can build a public-data pipeline (Google results, public LinkedIn profile snippets, company websites, podcast notes, public X profiles) for the cost of compute, proxy infrastructure, and the engineering time to keep the parsers from breaking when sites change their markup.
The math on scraping is much friendlier per contact: pennies, not dimes. The math on what those contacts are worth is brutal.
Three things break:
- Email verification rate collapses. A paid Apollo email is verified to ~95% deliverability accuracy. A scraped email (guessed from a name + domain via
firstname.lastname@,f.lastname@, etc.) verifies at 40-70% depending on company size and how much pattern matching you do. Anything below ~95% verification and your bulk sender reputation tanks. - Role accuracy drifts immediately. Static scraped data captures a snapshot. People change roles. The "VP of Sales" you scraped six months ago is now CRO at a different company, and the person you're emailing has just been promoted out of buying authority. Paid providers maintain decay-monitoring pipelines that catch this. Scrapers don't.
- Intent signal is absent. This is the silent killer. Paid enrichment can tell you "this person is VP of Sales at a company that just hit Series B." It cannot tell you "this person is in market for what you sell right now." Scraped data is even worse: it's the same static fact set, just less verified.
The deliverability collapse is the most quantifiable problem. If your bounce rate exceeds ~5%, Gmail and Outlook deliverability degrades; over ~10%, your sender domain gets flagged and reaches inbox at single-digit rates. An autonomous agent sending from polluted scraped lists doesn't fail loudly, it fails quietly, by being filtered to spam folders nobody checks. The agent reports "1,099 emails sent today" and the dashboard looks healthy, and zero of them reach an inbox.
This is constraint A and constraint C kissing. Scraped data triggers deliverability collapse, which triggers reliance on paid acquisition for distribution, which loops back to the unit-economics problem from post 2.
Why "smarter AI" doesn't break the wall
The intuitive response from someone who believes in the autonomous thesis is: "more powerful models will fix this. The agent will get better at qualifying prospects, deduping, verifying, intent-scoring."
That argument has a load-bearing assumption it doesn't earn. More inference cannot manufacture data that isn't present in the source. An LLM operating on scraped data can be more efficient at finding the good prospects within a low-quality pool, but it cannot raise the quality of the pool itself. The intent signal that wasn't captured isn't going to materialize because GPT-6 is reading the same scraped LinkedIn snippet.
Better models help around the edges: slightly better email-pattern guessing, slightly better role inference from public bio text. They don't change the fact that the highest-intent signal lives on platforms the agent can't access, in moments that don't show up in static enrichment data.
This is what makes constraint B the most structurally durable of the three. Inference gets cheaper, the agent gets smarter, the price-drop curve keeps cooking, and the data wall stays exactly where it is. In fact, as inference compresses toward zero, gated data becomes a larger share of the autonomous agent's competitive position, not a smaller one. Cheap intelligence operating on bad data still produces bad outcomes. The relative value of the data layer keeps rising.
The one bet that changes the equation
There's a single move that genuinely changes the math: stop buying static lists of people and start catching buyers in the moment they show intent. The highest-quality prospect data isn't a row in a database. It's a person posting a question, a complaint, or a "what do you use for X" right now, in public, before they've started shopping.
Not "scrape LinkedIn harder." Not "buy more Apollo credits." Instead: catch the moment buying intent surfaces in the open, and reach the buyer while the window is still warm.
Where does buying intent surface in public, in 2026?
- Reddit, Hacker News, and niche forums: founders and operators describe their pain in plain language before they ever run a search. "Ask HN: what tool do you use for X" is a buying signal in its clearest possible form.
- Bluesky and X: people vent about the exact problem you solve, in real time, in public.
- Public posts, podcast transcripts, GitHub issues: real-time intent layers that no enrichment vendor sells, because they aren't a list to sell.
The economic shape of this is a near-inverse of the paid-enrichment world. The signals are public and cheap to reach. The interpretation is the hard part: which comment is a genuine buying window vs. a casual mention, which thread is a serious pain vs. venting, and how you respond in a real human voice instead of the AI slop everyone is sick of. That's where the value is created, and it's exactly the part a list vendor can't sell you.
And here's the part the autonomous-agent thesis can't easily copy: this kind of outreach is high-intent and low-volume by structure. Blasting every mention at autonomous volume is how accounts get flagged and banned. Reaching the right buyer at the right moment, in your voice, with the right pacing, is the opposite of the "1,099 emails sent today" pattern. That's why it lands in inboxes instead of spam folders.
The move that closes constraint B doesn't look like "smarter AI on the same data." It looks like getting in front of the buyers who are already raising their hands, before your competitors even know they exist.
The honest read
The data wall isn't getting weaker. Apollo isn't lowering prices. ZoomInfo isn't opening their database. LinkedIn isn't loosening enforcement. If anything, all of them are getting more aggressive about gating as AI-driven scraping makes their data more contested.
The autonomous-agent platforms that try to solve constraint B by buying more Apollo credits will discover the unit economics don't close. The ones that try to solve it by scraping harder will discover deliverability collapses. The ones that try to solve it by waiting for "AI to get smarter" will discover smarter AI on bad data is still bad outcomes.
Full autonomy is the wrong target at SMB economics. The move that actually wins is autopilot for the grind that pays off: find the buyers already asking for what you sell, draft the reply in your voice, and pace the sends so your accounts stay healthy and land in inboxes. That puts leads and customers in front of you, each with a receipt for the conversation that produced it, instead of burning revenue on data that was never going to close the math. The founder sets the bar and the voice; a judge holds that bar on every single draft. That's a business that closes the math today, not in 24 months.
Next in the series: constraint C, the distribution problem. Why "distribution is bought, full stop" is the most honest sentence in the autonomous-agent industry, why the cold channels are closing, and why organic distribution requires the one thing the autonomous thesis structurally cannot manufacture.
This is post 3 of a 5-part series on the AI-runs-your-company thesis. Start with post 1 for the full audit, or read post 2 for the inference math.