

AI agents can reason through almost anything. Ask one to actually get something, like live data, a voice, a map or a flight price, and it hits the same wall every time. Someone has to open an account, paste a key and put a card on file.
At SF Tech Week I tried to make that wall visible. My talk, "AI agents are the new customers," was hosted by Apify at the European Startup Embassy. I ran one prompt in two terminals, side by side, live. Same model, same prompt. One agent had the Nevermined Catalog and a budget. The other was told not to use it.
The prompt was greedy on purpose. It asked for a web app of the best AI-agent and agentic-commerce events in the next 60 days, with a voiced news brief, maps, weather, flights, translations and a public link at the end. One terminal finished with a working product and a receipt for $1.82. The other finished with a well-written page full of placeholders, each one naming an account a human would have to open.
Both agents were equally capable. The only difference was what each one was allowed to buy. This post walks through what one prompt bought, what it cost, how the catalog decides what to buy, and why I think this matters far more for businesses than for consumers.
The agent with the catalog came back with Agent Events Radar, and it bought every part of it. It used 22 catalog services, made 75 paid calls and spent $1.82 in total. It needed no API keys and no signups. The agent hosted the page on StableUpload, also bought from the catalog, which keeps it online for six months. I wanted it to last longer, so I also pinned the page to IPFS through Pinata, from the same catalog, for about a cent (IPFS copy).

Here is what that $1.82 bought:
The agent never saw a list of service names. It described each need in plain words, like "find events in Europe" or "read this page aloud," and the catalog picked a service. When a pick failed, it moved to the next one on the shortlist. The stock-quote service it tried first returned an error, so it switched providers. The failed call was never settled and cost nothing.
The ledger also shows where the money actually went. Most calls are tiny. The median call cost about one cent, and 33 of the 75 settled payments were under a cent. Screenshots and translations made up half the bill, because those are the jobs with real compute behind them. The one call that bought nothing useful was 18 cents on a trending-topics search, which came back with an answer the agent couldn't verify, so the page doesn't show it.

That receipt only means something next to what the other terminal produced. The agent without the catalog built this page, and I want to be fair to it, because in places it did better. It found four cited news stories to our two. Its events covered five countries to our three. It pulled stock quotes from a free Yahoo Finance endpoint and exchange rates from the ECB.

What stands out is everything it left empty, and why. It didn't run out of intelligence. It ran out of things it could buy. It used five services, spent nothing and listed sixteen sections it couldn't fill. For each gap it named the account a human would have had to open first:
The catalog didn't do everything either, and the honest version of this story says so:
Look again at that last column in the grid. Every gap names an account, and that is the real cost the second agent ran into. It was never the price of the calls. It was the setup. One real job touches about nine services. Each one wants its own account, its own API key, a card on file and an invoice to reconcile. That's 36 chores before the first result, and an agent can't do any of them, so the job goes back to a human.

The catalog collapses those 36 chores into one account, one budget and one ledger, with no per-service keys. You top up once with a normal card, and nothing is subscribed or pre-committed. From there, the pricing comes down to four rules:

The Radar run shows those rules working. I gave the agent a $6 budget with a 30-cent limit on any single call, and told it to set its own maximum on every call. When one service quoted 18 cents against a 10-cent maximum, the router refused it before anything was paid. The agent had to decide to raise the limit and ask again. The budget ledger tracks spend to a ten-thousandth of a cent, so the $1.82 on the page is the same $1.82 on the ledger.
Platforms get one more layer, and it's the one I spend most of my time on. A development platform, a model router or an API provider can bring the catalog to its own users. You buy at the published price, add your own margin and put it on the bill your users already pay. You can wrap it, pass it through, blend it into your credits or bundle it into a subscription. A larger commitment buys a better wholesale rate or a bigger allotment. Your users see your product, and Nevermined stays in the background.

Cheap calls only matter if the agent buys the right thing, which raises the obvious question. How did the agent end up with that screenshot service and that voice? The honest answer is that it didn't choose. It asked the shelf in plain words, and the catalog's Auto mode chose. Auto scores the shelf against the request, returns one listing and pays for the call. No model reads the listings, and every pick comes back with the shortlist and the reasons, so you can see why the winner won.
It starts with a gate. Before anything is ranked, the router drops every listing it can't actually pay. In the Radar run, a search for nearby restaurants came back "no fundable match" at no charge, and the agent moved on instead of guessing. Every listing also gets a real, unpaid request every hour. The router reads back the payment challenge, nothing is paid, and a listing stays on the shelf only by proving it can still take money.
What passes the gate is ranked with two published formulas. The first is a quality score between 0 and 1, built from what the router observes when real calls go through:
The second is the rank score, the public sort key, which adjusts quality for two things we know about the listing:

Read it from the bottom up. What the router observes on real calls becomes quality, and quality is adjusted only for curation and for whether a paid test call has passed.
Here is what every term means:
When an agent asks for something, Auto blends relevance with rank: relevance × (1 + 0.25 × q), where q is the listing's rank score divided by the best in that set. Relevance comes first, and quality decides between listings that fit.
Two details matter for trust. Traction carries the most weight, at 45%, and only the party routing the payments can see it, because every settlement feeds the next ranking. And the formula used to include a multiplier for paid organization tier. We took it out, because it wasn't disclosed and it contradicted the catalog's neutrality. I'll also be upfront that the shelf is young: for most listings, traction and settled-price signals are still near zero, so today the rank leans on liveness and the quoted price. It gets sharper as more agents pay.
An events page is a fun demo, but the pattern underneath it is what matters. A person describes an outcome, and the agent assembles it from a dozen paid services under one budget. Once you see it that way, the use cases run from one person's evening all the way up to a company's fleet of agents.
It starts with individuals, because that's where most people first meet an agent. Paste in a suspicious text and the agent checks the sender's domain, the site's server and whether the email address is real. Ask whether you can afford a neighborhood and it pulls the median rent and your commute. Each of these is a few cents of search, data and media that nobody would open five accounts to get, and I tested both this week (more below).
We've seen this outside my demo too. Aitor on our team connected the catalog to Claude, ChatGPT and LangSmith Fleet with one URL and one budget he approved, $5 for seven days, and no API keys. Claude compared browser privacy claims from the vendors' live pages. ChatGPT planned a weekend in Lisbon. A LangSmith agent wrote a company brief from a live website. Five vendors got paid, a few cents in total. His line stuck with me: the agent only pays for data it can't reach, and the thinking is free.
The same pattern gets more valuable once a business depends on it. Sales, marketing, finance and legal teams already pay for research spread across tools they buy one at a time. An agent with a budget can do that research on demand.
I didn't want to list those jobs from imagination, so this week I ran twelve of them through the catalog with real money. About 100 paid calls went out across 50 services, and 77 of them came back with something useful. Here's what each job cost and what I found:
The honest summary is that the shelf is broad and still uneven. Most jobs worked for pennies. A few services charged for empty results or failed in ways an agent can't learn from yet, and a couple of listings promised more than they delivered. That's exactly why Auto ranks on what real calls return: every one of these payments feeds the next ranking.
From there, the step to the enterprise is about control, not capability. A company running hundreds of internal agents can't hand each one a corporate card and a drawer of API keys. Instead, every agent gets its own budget with a cap, an expiry and a per-call limit, and every call lands on one ledger. Finance sees the merchant and the amount on every call, and a budget is revoked in one click. Security sees no vendor keys sitting in anyone's code. As Aitor put it, with an API key you find out at the end of the month, and good luck working out which agent did it.
Platforms are where this compounds. A development platform can give every agent built on it a tool that can pay. A model router can offer a services lane next to the models it already routes. An API provider can list its own service where every agent's Auto mode can find it, and buy everything else on the same account. In each case the platform keeps the customer relationship and sets its own margin.
That path from one person to a fleet of agents is also how the money moves. The consumer forecasts get the headlines, but the B2B one is an order of magnitude larger. Each firm also measures something different, so I'd rather show you the definitions than pick the biggest number:
The consumer numbers range from "influenced" to "completed," which is why they're spread so wide. McKinsey is explicit that its figure leaves out services and the B2B marketplace entirely. Gartner's $15 trillion is a strategic prediction rather than a sized market model, but it's the clearest signal of where the volume is headed: business buying, done by agents on behalf of teams.
I also want to be careful here, because the gap between forecast and today is real. Gartner itself expects more than 40% of agentic AI projects to be canceled by the end of 2027 (via OODA Loop). And the payments agents measurably make on their own today add up to tens of millions of dollars, not trillions. That gap is the opportunity, and it's mostly a plumbing problem, the same one the second terminal ran into.
There's a narrower market inside all of this that gets less attention, and it's the one the catalog serves. Before an agent buys a pair of shoes or a software seat, it buys inputs: a search, a page read, a data lookup, a voice, an image. That's what our Radar agent spent its $1.82 on. Today those inputs are sold through API keys and monthly plans built for human developers. Postman's 2025 State of the API found that 65% of organizations already earn revenue from their APIs, yet only 24% of developers design those APIs with AI agents in mind. Nobody has sized "agents paying per API call" as its own market yet. I think that's because it barely existed until agents had a way to pay.
Forecasts are one way to think about this. Watching your own agent do it is better, so we're giving builders up to $50 in catalog credits. The shelf has more than 160 services across 13 categories: search, scraping, enrichment, filings, market data, media and more.

Here's how the credits work. You get $10 to start, which is about 500 calls, or five builds the size of the Radar page. Then tell us what you loved or hated. A person reads it and tops you up to $50. There's no card and no demo call.

Then connect your agent from the get-started page. The fastest path takes one command in Claude Code. Your browser opens, you log in and approve a spending cap, and there's no API key to generate or paste:

If you'd rather use an API key, for scripts or a harness without OAuth, the page walks you through four steps. You log in with Google, GitHub or email, then top up your balance; with claimed credits, that's already done. Next, create one key under Dashboard, API keys. Last, call any service by prompt, skill, MCP or cURL. It works the same in Claude Code, Codex, Cursor, OpenCode, Gemini CLI or any assistant that speaks MCP.
Then ask for an outcome, not a service. Here's a good first prompt:
"Using the Nevermined catalog, find the latest funding news for these five companies, read each company's homepage, and give me a one-page brief with sources. Set a $1 budget, cap each call at 10 cents, and show me a receipt of what you spent."
Watch what comes back. The agent searches the shelf, lets Auto pick a service for each step, pays for each call inside your budget and hands you the receipt. If you want to go further, paste in the Agent Events Radar prompt and see what your agent builds for a couple of dollars. The best builds may get listed in the catalog free for five months.
When I started this, the question was whether an agent could do something significant on its own if we gave it a shelf and a budget. The answer cost $1.82. Even the narration came off the same shelf, for under 50 cents including retakes. I'd love to see what you build with yours: claim your credits and send it to me.
It felt wrong to write about receipts without showing this one. Everything bought for this post went through the same catalog, on the same budget, and the total was about $5.60:
Everything else came from free tools and our own material, so it isn't on the bill. Claude did the research and the drafting, and its own usage isn't counted above. The market research used Claude's built-in web search and page reader, not the catalog. The screenshots of both pages, and the images that compare them, came from a local browser. I drew the scoring diagram, the call flow and the spend chart from the run's ledger and our published scoring page. The source material was our AI Services Catalog for Platforms deck, the Services Router pitch, the catalog's scoring page, the get-started and claim pages, and Aitor's post and video.

See Nevermined
in Action
Real-time payments, flexible pricing, and outcome-based monetization—all in one platform.