When to Use a Scraping API vs Building Your Own Scraper
Pieter Levels said he would keep paying $99/mo, then replaced a $249/mo plan with a $1/mo scraper five days later. Here is what the credit math actually says.
On September 4, 2026, Pieter Levels wrote that he'd spent weeks trying to build his own ScrapingBee and failed, so he'd keep paying for it. Five days later he wrote that he'd replaced a $249/mo plan with a $1/mo scraper on his own VPS.
Same person, same problem, opposite conclusion, five days apart. That's a rare thing to get in public, and it's a better guide to this decision than any vendor comparison, because both posts are honest about what broke.
So where's the actual line? Roughly here. If you're scraping one or two known sites at low volume and you can afford to go slowly, build it. If you need breadth, speed, or reliability you don't have to babysit, pay for the API. What decides it isn't your code. It's whether you can get residential IPs that Google hasn't already flagged.
Every price below came from the vendor's own pricing page this week.
Quick Decision Table
| Your situation | Use this |
|---|---|
| One or two known sites, under ~3,000 requests/day, pacing is fine | Self-hosted Patchright on a VPS |
| Google SERP work above ~2,000 requests/day | ScrapingBee Startup or Business |
| Many different sites, unpredictable anti-bot | ScrapingBee or ScraperAPI |
| Cost per request matters more than convenience | Zyte, from $0.06 per 1,000 |
| Mostly plain HTML, no JS rendering needed | Firecrawl, 1 credit per page |
| You can't tolerate a scraper silently breaking | Any managed API |
What Actually Happened in Those Five Days?
In the September 4 post, Levels said 80% of his scraping is Google SERP results for Hotelist, and that residential proxies, CAPTCHA-solving services and OpenSERP all failed on "thousands of queries". His conclusion: "So I will keep paying them $99/mo". The same day, replying to someone on X, he added "I am happy with Scrapingbee" and said the point was that he wanted to replace it with his own.
Then the September 9 post landed, crediting help from Javi Lopez and, in his words, "his Spanish scraping friends". What changed was the browser, not the ambition:
- OpenSERP failed because it drives a vanilla headless Chrome. Google fingerprints that instantly. He quotes his AI's measurement: "plain curl got 7% blocked but OpenSERP's headless browser got 100% blocked"
- IPRoyal residential proxies failed because the IPs were already burned. Google "just blocked about 60% of them already"
- What worked was a real Chrome, headful, with a persistent context, on his VPS, plus "slow pacing like every 30 seconds or so"
- He got better proxy recommendations he won't share, "because then they stop working"
Result: about 90% success, "which is similar or better than Scrapingbee", for around $1/mo in bandwidth.
Worth noting he published a third post the same day listing roughly $25,000/mo of SaaS he'd replaced. Its scraping line reads "Scrapingbee/SerpAPI -> My own $1/mo scraper with Patchright". So the tool is Patchright, a stealth-patched Playwright fork, not plain Playwright. And the thing replaced was two services, not one.
About that $99 and that $249
The two posts don't agree. On September 4 it's $99/mo. On September 9 it's $249/mo, twice. Both are real adjacent ScrapingBee tiers, Startup and Business, and nothing in between the posts mentions upgrading.
The third post's "Scrapingbee/SerpAPI" line suggests the $249 might be a combined bill, and a ScrapingBee Startup at $99 plus a SerpAPI Production at $150 sums to exactly $249. That's arithmetic that fits, not something he says. I'm flagging the gap rather than picking the bigger number, because the bigger number is the one that makes the story sound better.
Why Is a Credit Not a Request?
ScrapingBee bills credits, not requests, and the multiplier swings 75x depending on what you ask for. From their credit system documentation, updated August 2026:
| Request type | Credits |
|---|---|
| Classic proxy, no JS rendering | 1 |
| Classic proxy with JS rendering (the default) | 5 |
| Premium proxy, no JS | 10 |
| Premium proxy with JS | 25 |
| Stealth proxy | 75 |
| Google URL request, classic or premium | 15 |
| Google URL request, stealth proxy | 75 |
So "3,000,000 credits" on the $249 Business tier is 3 million requests only if you scrape plain HTML with no rendering. For Google SERP work at 15 credits a request, it's 200,000 requests a month. About 6,700 a day.
Now put his scraper next to that. One request every 30 seconds is 2,880 a day.
Those are the rates for classic or premium proxies, which is what a Google URL request uses by default. Push a target onto stealth proxies and it costs 75, which divides every number below by five.
| Option | Google requests/day | Cost/mo | Cost per 1,000 |
|---|---|---|---|
| ScrapingBee Hobby | 167 | $19 | $3.80 |
| ScrapingBee Freelance | 556 | $49 | $2.94 |
| ScrapingBee Startup | 2,222 | $99 | $1.49 |
| ScrapingBee Business | 6,667 | $249 | $1.25 |
| Self-hosted, paced at 30s | 2,880 | ~$1 plus VPS | varies |
Unit costs rounded to the cent.
That reframes the $249 to $1 headline. His scraper isn't running at a fraction of what he was paying for. At that pacing it does slightly more Google requests per day than the $99 Startup tier allows, and about 43% of the Business tier. The saving is real. The throughput gap is much smaller than the price gap makes it sound, which is what makes this worth copying rather than dismissing.
When Should You Build Your Own?
Four things need to be true at once.
Your target list is short and known. One or two sites you understand, where you can tune for their specific defenses. Breadth is what kills self-hosted scrapers, because every new domain is a new anti-bot puzzle.
Your volume fits inside slow pacing. Pacing is the trick that replaces proxy rotation. At one request per 30 seconds you get roughly 2,880 a day per IP, and that's enough for a lot of indie projects. Need 50,000 a day? You're back to buying proxy pools, and then you're paying for the thing the API was selling.
You can get clean residential IPs. This is the hard gate, and it isn't a code problem. IPRoyal starts at $7/GB on a subscription, or $7.35 pay as you go, with no commitment either way. Sounds fine until 60% of the IPs turn out to be pre-flagged. Levels solved it through private recommendations he won't publish, and he's right that publishing them would burn them. You may not have that network.
A silent failure is survivable. A managed API failing is someone else's pager. Your scraper failing is a quiet gap in your data that you find out about later.
When should you NOT build your own?
If your scraping is load-bearing for a paid product, buy it. A 90% success rate that was three days old when it was written up is not a reliability figure yet, and Levels didn't claim it was.
If you need more than a handful of sites, buy it. If you need real throughput, buy it, because the cost moves from your VPS bill to your proxy bill and the API's volume pricing starts winning. And if you don't already have a line on working residential proxies, buy it, because that's the part no amount of AI assistance solves.
What Do the APIs Actually Cost in 2026?
Prices verified this week from vendor pricing pages.
ScrapingBee: Hobby $19 (75k credits), Freelance $49 (250k), Startup $99 (1M), Business $249 (3M), Business+ $599 (8M). 1,000 free credits, no card. Concurrency runs 25 to 400 by tier. Prices exclude VAT. Worth knowing: ScrapingBee joined Oxylabs' group on January 19, 2026, said plans stay the same, and cut Google API calls from 25 credits to 15. So the acquisition made Google scraping cheaper rather than dearer, which is the opposite of the usual post-acquisition story.
ScraperAPI: Hobby $49 (100k credits), Startup $149 (1M), Business $299 (3M), Professional $975 (10.5M). About 10% off annually. 1,000 free credits. Note the entry point is $49 for 100k against ScrapingBee's $19 for 75k.
Zyte: usage-based, from $0.06 per 1,000 successful responses, priced by an automatic 1 to 5 difficulty tier. Pay-as-you-go HTTP runs $0.13 to $1.27 per 1,000, browser rendering $1.01 to $16.08. A monthly commitment buys a better rate. $5 free credit. The cheapest option if your targets are easy, and the most expensive if they're hard.
Firecrawl: Free 1,000 credits, Hobby $16/mo, Standard $83/mo, Growth $333/mo, Scale $599/mo, all on annual billing. Credits are refreshingly literal: 1 credit per page, search costs 2 per 10 results. No JS-rendering surcharge, which is a real differentiator against ScrapingBee's 5x default.
And the self-hosted side?
The old advice to grab a €3 Hetzner box is dead twice over. The cost-optimized CX line shows as unavailable, and Hetzner raised shared vCPU prices on June 15, 2026. That increase applies to new orders and rescales, so if you already run an older box you keep your old rate.
Headful Chrome needs real memory, so CPX12 (1 vCPU, 2GB, €11.49/mo) is marginal and will run out the moment you open a second context. CPX22 at 2 vCPU and 4GB, €19.49/mo, is the realistic floor. Want parallelism? CPX32 at €35.49. Those are EU location prices net of VAT, and Singapore runs about 35% higher.
So the honest self-hosted number isn't $1. At today's rate €19.49 is about $23, and a gigabyte of residential traffic from IPRoyal is $7, so budget roughly $23 to $30 a month depending on how much proxy bandwidth you actually burn. Against $99 for the tier it replaces, still a clear win. Just not a hundredfold one.
My Recommendation
Build it if you're scraping one or two sites you know well, under about 3,000 requests a day, and you already have residential proxies that work. That's a real and reachable setup, and Levels' numbers back it up.
Pay for ScrapingBee if Google SERP work is central and you need more than a couple of thousand requests a day. The Startup tier at $99 for 66,666 Google requests a month is hard to beat once you price your own time in. Pay for Zyte instead if your targets are easy and volume is high, since $0.06 per 1,000 undercuts everyone. Use Firecrawl if you're mostly pulling plain pages and the 5x JS multiplier annoys you.
And if you're somewhere in the middle, run both. Self-host the one site that's 80% of your volume, and keep a small API plan for the long tail. That's the setup most people should land on, and nobody sells it because no vendor makes money recommending half a subscription.
flowchart TD
A{Clean residential proxies available?} -- no --> B[Pay for an API]
A -- yes --> C{More than 2 or 3 target sites?}
C -- yes --> B
C -- no --> D{Over ~3,000 requests/day?}
D -- yes --> B
D -- no --> E{Can a silent failure wait a day?}
E -- no --> B
E -- yes --> F[Self-host Patchright on a VPS]
We ranked the managed options in the AI web scraping tools roundup, and the browser automation trade-offs in Playwright vs Cypress vs Selenium. If you're picking a box for the self-hosted route, when to use Vercel vs Railway vs Hetzner covers why a plain VPS keeps winning this kind of job.
The decision isn't really build versus buy. It's whether you have access to clean IPs, because that's the only input here you can't write yourself.
Found a better option? Let me know on Twitter @devtoolpicks.
Frequently Asked Questions
Is it cheaper to build your own scraper than use ScrapingBee?
At low volume on a small number of sites, yes. A Hetzner CPX22 runs about 19.49 euros a month and residential proxy traffic for text scraping costs a few dollars, against 99 dollars for ScrapingBees Startup tier. The saving disappears once you need many sites or high throughput, because you then pay for proxy quality and maintenance instead of code.
How many credits does a Google search cost on ScrapingBee?
A Google URL request costs 15 API credits on classic or premium proxies, and 75 on stealth proxies. That matters because credits are not requests. The 249 dollar Business tier includes 3,000,000 credits, which is 200,000 Google requests a month, or about 6,700 a day. A plain request with no JavaScript rendering costs 1 credit.
Why did OpenSERP not work for Google scraping?
OpenSERP drives a vanilla headless Chrome, which Google detects immediately. Pieter Levels reported that his AI measured plain curl at 7 percent blocked against OpenSERPs headless browser at 100 percent blocked. Headless Chrome leaks automation signals that Google fingerprints. Tools like Patchright exist specifically to patch those leaks out of Playwright.
Do residential proxies fix scraper blocking?
Only if the IPs are clean, and many are not. Levels tried IPRoyal residential proxies and found Google had already flagged roughly 60 percent of them as scrapers. IPRoyal starts at 7 dollars per GB. Proxy reputation is the part you cannot fix with better code, and the providers that work tend to stop working once they get named publicly.
What is Patchright and why use it instead of Playwright?
Patchright is an Apache 2.0 licensed fork of Playwright that removes automation fingerprints, and it works as a drop-in replacement. It avoids the Runtime.enable leak, disables the Console API, and strips flags like enable-automation. It patches Chromium browsers only, so Firefox and WebKit are not supported. It has about 4,600 GitHub stars.
Get honest tool comparisons in your inbox
Join 50+ indie hackers and solo developers who get new comparisons, pricing changes, and tool picks. No spam. Unsubscribe anytime.
Related Articles
Best AI Web Scraping Tools for Indie Hackers in 2026
Four different ways to pay for web data in 2026, from flat credits to pure meter...
When to Use LiteLLM Self-Hosted vs a Managed AI Gateway
ChatGPT, Claude and Grok all went down on the same morning. A gateway is how you...
Best Cron and Scheduled Job Monitoring Tools for Indie Hackers in 2026
Your nightly backup stopped running three weeks ago and nothing alerted you, bec...