TL;DR
- Firecrawl is our default for turning known company pages into agent-ready text. Scrape selected URLs; crawl narrowly when useful pages still need locating.
- Apify is strongest when the workflow needs a reusable Actor or custom extraction code, with scheduling, stored results and control over HTTP versus browser execution.
- Zyte fits developers who want explicit HTTP/browser retrieval and schema-based extraction, priced by the target site's difficulty and selected features.
- Bright Data fits supported structured sources such as company profiles and job boards. Its per-record Scraper API price should not be treated as the universal cost of extracting arbitrary company websites.
- On 3,000 pages, operator review can cost far more than retrieval. Keep a source passage and return missing fields as missing rather than filling a schema with plausible guesses.
Choose the retrieval depth
Separate discovery, retrieval and interpretation. Search locates a candidate URL. Retrieval obtains the page. Extraction turns its content into a field or clean text that the research agent can use. This guide focuses on known company, careers and product pages, with structured profile/job sources as a separate route.
A static company description may need only HTTP and a parser. A careers page that loads jobs with JavaScript needs rendering. A page requiring a click to reveal a role needs an action. These are different workloads, even when the final field is simply “recent engineering hiring.”
Start with three selected pages per account, rather than crawling every link. Restrict paths and page count, discard duplicate canonical URLs and keep the source date separate from fetch time. A page retrieved today can still describe an old event. The execution layer is covered in running GTM from a coding agent.
The four providers compared
| Provider | Output and execution | Current pricing basis | Best fit | Main non-fit |
|---|---|---|---|---|
| Firecrawl | Markdown/HTML, custom JSON, batch scrape and scoped crawl; rendering and actions | Hobby $19/month, 5,000 credits; base page one credit, JSON adds four | Agent-ready content from known URLs | Treating every page as a five-credit JSON job when text is sufficient |
| Apify | Hosted Actors, custom code and datasets; HTTP or browser crawler modes | Starter $19/month includes $19 usage; $0.20 per CU, with other resource/Actor charges | Repeated custom extraction and scheduled jobs | Assuming one universal per-page tariff across Actors |
| Zyte | HTTP body or rendered HTML; automatic types and custom attributes | PAYG HTTP $0.13–$1.27/1,000; browser $1.01–$16.08/1,000, plus selected features | Controlled retrieval and extraction with a developer-owned pipeline | Applying the lowest HTTP tier to all dynamic sites |
| Bright Data | Supported-site Scraper APIs return JSON/CSV; sync or async delivery | Scraper API PAYG $1.50/1,000 delivered records | Standardised company/profile/job source records | Equating a supported-source record with any arbitrary webpage |
Prices were checked on 30 September 2026, before tax. Firecrawl's monthly option is used in the budget; its $16 Hobby display is the annual billing equivalent. Firecrawl pricing, Apify pricing, Zyte pricing, Bright Data Scraper API pricing.
Product capabilities and fit
1. Firecrawl: best default for agent-ready account pages
Firecrawl is a focused option for retrieving a company page as clean Markdown or HTML, then passing selected evidence into a research agent. JSON mode accepts a custom schema, making it useful for fields such as hiring location, product category or a dated announcement. Its scrape interface also supports actions and location options. Scrape documentation.
Use batch scrape when URLs are already known. Use Crawl when the agent needs to discover pages within a site, with path filters, depth and link controls. Set an explicit page limit: the documented default is 10,000, a poor starting point for a three-page account brief. Crawl documentation.
Freshness has a concrete control. Scrape's default cache window is two days; maxAge: 0 requests a new fetch. Cached results still consume a base credit, so the application's own saved evidence can reduce repeat requests more effectively than expecting provider caching to make them free. Use a short window for open roles and a longer one for stable product descriptions. Cache behaviour.
Basic scraping consumes one credit per page; JSON adds four. A scrape with no result is uncharged, but a returned 403 or 404 page costs a base credit. That is a reason to validate the returned page before asking a model to analyse it. Current credit rules.
Non-fit: complex site-specific workflows whose navigation and field rules are already custom code may be easier to operate as an Apify Actor. Firecrawl also offers search; describing it as incapable of discovery would be inaccurate. The practical choice here is its known-page scrape/crawl path.
2. Apify: best for reusable extraction jobs and custom logic
Apify packages scraping and automation into Actors that can be invoked through APIs and produce stored datasets. Its first-party Website Content Crawler cleans page content into text or Markdown and exposes output as JSON/CSV. It accepts start URLs, crawl boundaries and page/depth controls. Website Content Crawler.
The current crawler's adaptive mode switches between raw HTTP for static pages and Firefox/Playwright for dynamic ones. Raw HTTP is the lighter option when it captures the required content; browser mode can wait, scroll and expand elements. The older JSDOM and Chrome/Playwright modes are marked deprecated in this Actor, so they should not be the default implementation recommendation. Crawler modes.
Apify is useful when an engineer needs to encode stable selectors, traverse pagination or maintain different adapters for several careers systems. Instead of asking a model to infer the same page layout on every request, the Actor can emit consistent records and leave only interpretation to the model.
Starter's $19 monthly payment is prepaid usage, not $19 added to every resource bill. A compute unit is one GB of RAM used for one hour, priced at $0.20 on Starter. Residential proxy traffic is $8/GB there. Storage, transfer and a chosen Actor's own pricing can also matter; Store Actors do not all use the same charging model. Platform pricing.
Non-fit: a team needing a one-call Markdown response from a handful of pages may find Firecrawl simpler. An Actor still needs a maintainer when a target site's structure changes; a marketplace listing is not a permanent extraction guarantee.
3. Zyte: best for explicit retrieval and extraction control
Zyte exposes the distinction between httpResponseBody and browserHtml. The former retrieves the HTTP body; the latter returns rendered DOM content. Automatic extraction includes types such as jobPosting and pageContent, with the extraction source configurable. Select the source deliberately rather than paying for a browser when HTTP contains the needed evidence. API reference.
Custom attributes add an LLM-generated schema output to a standard extraction field. For arbitrary company pages, pageContent is the general-purpose basis; for an individual role, the job-posting type is more appropriate. Custom fields are useful when careers pages vary, but the values still need source grounding. Custom attributes.
PAYG has five site tiers for each retrieval type. HTTP ranges from $0.13 to $1.27 per thousand responses, while browser rendering ranges from $1.01 to $16.08. These are retrieval bases, rather than complete structured-extraction prices. Automatic extraction adds $0.0004–$0.0016 per data type; custom attributes have token-based costs, and actions or screenshots have separate meters. Commercial rates, Feature charges.
Zyte charges successful responses, but a page not matching the selected extraction type need not produce an API failure. A successful response and a usable hiring record are different outcomes. Error handling.
Non-fit: a nontechnical user expecting a finished account-research product. Zyte reduces retrieval infrastructure work, while the team still owns URL selection, parsers or schema, qualification logic and delivery to the CRM.
4. Bright Data: best for supported structured source records
Bright Data's Scraper API library provides pre-built source-specific extraction. For account research, relevant examples include company information and job listings. The interface supports a synchronous response for a small request, asynchronous jobs for batches, and API, webhook or storage delivery. Scraper API overview.
Buy this route when the source and fields already match a supported scraper. Receiving a structured company or role record can eliminate a separate parser. Scraper Studio is a different route for a custom target; browser and other web-access products also have their own pricing.
The directly applicable PAYG card lists $1.50 per thousand records, with parsing and retrieval infrastructure included and failed deliveries uncharged. Its free tier lists 5,000 monthly records. The navigation's lower “starts from” headline is a volume starting rate, rather than the PAYG tariff used here. Scraper API pricing.
Non-fit: three arbitrary URLs from every prospect's website under the assumption that they produce three supported-source records. The separate Crawl API page lists a $1.50/1,000-request basis, but requests, crawled pages and delivered records should not be substituted for each other. Choose the matching endpoint and output model, rather than applying a single Bright Data price to its whole catalogue. Crawl pricing basis.
Account research output and retries
For a public careers page, aim for a compact evidence object:
{
"account_id": "A-204",
"company_domain": "northstar-controls.example",
"source_url": "https://northstar-controls.example/careers/",
"fetched_at": "2026-09-30T08:00:00Z",
"signal_type": "engineering_hiring",
"role_title": "Controls Engineer",
"location": null,
"posted_at": null,
"evidence_excerpt": "We are recruiting a Controls Engineer.",
"status": "evidence_found_date_unknown"
}
This is a fictional output contract, not a response reported from a provider trial. It leaves location and date empty because the example source does not establish them. The fetch timestamp is not converted into a posting date. A downstream agent can write a useful note saying the careers page lists the role without inventing a recent hiring event.
Retry transient failures with a bounded attempt count. Escalate from HTTP to rendering only when useful content is missing, rather than rendering every page. Store the final URL, response status, extractor version and a content hash. A consent page or empty application shell should produce a retrieval exception, not a fabricated company description.
Use a stable job ID for asynchronous batches. If the polling request fails, retrieve the existing job's output rather than blindly resubmitting the entire crawl. On Apify, retries can consume compute; on per-page or successful-response systems, another delivered result can consume another unit. “No charge for failures” does not make every application retry free.
Worked budget for 3000 pages
Assume 1,000 accounts with three known URLs each: 2,400 static pages and 600 rendered careers pages. No full-site crawl, screenshots or PDF documents are included. These are workload assumptions, not measured success rates or runtimes.
| Provider route | Unit calculation | Supported budget |
|---|---|---|
| Firecrawl, Markdown on all pages | 3,000 base credits | $19 Hobby monthly subscription, within 5,000 credits |
| Firecrawl, JSON on 600 careers pages only | 3,000 + 600 × 4 = 5,400 credits | $19 + one $5 increment of 1,000 credits = $24; 600 extra credits remain |
| Apify, usage-priced custom/first-party job | Assumed 50 CU × $0.20 + 2 GB residential proxy × $8 = $26 | $26 for those metered resources, using the $19 prepaid allowance and $7 extra; storage/transfer or a selected Actor's other fees separate |
| Zyte, illustrative Tier 2 mix | 2,400 HTTP × $0.23/1,000 + 600 browser × $2.01/1,000 | $1.758 retrieval base; extraction and other selected features additional |
| Bright Data, alternative supported-source workload | 3,000 delivered records × $1.50/1,000 | $4.50 gross PAYG basis before free allocation; applies to supported records rather than the arbitrary-page set |
The comparison uses recurring paid bases and excludes free or introductory credits from the arithmetic. Apify's CU and bandwidth figures are explicit estimates. Zyte's example assumes Tier 2 targets: at Tier 5 rates the same static/browser mix has a $12.696 retrieval base. That range shows the sensitivity without pretending every target has the cheapest price. Apify resource rates, Zyte tier rates.
Adding one Zyte automatic-extraction type to all 3,000 results adds $1.20–$4.80 before any applicable discount; custom attributes and actions remain additional. Request the structured fields only where they remove work. Zyte feature pricing.
On Firecrawl, asking for JSON on all 3,000 pages would use 15,000 credits. Hobby's 5,000 included credits plus ten $5 increments gives a $69 monthly basis. That may still be worthwhile, but the selective $24 design is a clearer starting point when only careers pages need structured fields. Firecrawl credit rules.
The larger cost is review. If 300 pages need two minutes of checking, that is 600 minutes, or ten hours. At an assumed $50/hour, review costs $500. Reducing exceptions to 150 pages saves five hours, worth $250. Against that saving, a few dollars of retrieval-price difference is minor. Compare providers on usable source-backed fields per account, not raw successful requests.
Choose the provider by workload
Use Firecrawl for a new agent that needs selected company-page text quickly, with JSON reserved for fields that justify its credits. Use Apify when extraction is an operated job with custom code, recurring schedules or adapters that the team wants to own. Use Zyte when developers need deliberate HTTP/browser routing and native extraction types, accepting site-tier pricing as the cost model.
Use Bright Data when a supported company/profile/job source gives the records the workflow needs. Keep arbitrary-page retrieval as a separate path. An existing scraper that returns useful structured evidence is more valuable than swapping providers for a lower headline unit price.
Tie extracted fields to the qualification criteria in the signal-based outbound playbook. Work with public or authorised sources and keep the minimum evidence needed for the buying signal. The pipeline is complete when the agent can explain its account note from the retrieved passage, rather than merely populate every field.
FAQ
Does every account page need a browser?
No. Use HTTP when it contains the relevant content. Render pages whose useful content depends on JavaScript or an interaction.
Is Firecrawl only a scraper?
No. It also provides search and other web-data capabilities. For known URLs, scrape or a bounded crawl is the appropriate route.
Does Apify charge one price per page?
No. Actors and platform resources have different pricing models. A compute unit measures one GB-hour, while proxies and other resources can add separate charges.
Is Zyte's lowest price an all-in extraction rate?
No. It is a site-tier HTTP retrieval rate. Rendering, automatic extraction, custom attributes and actions have different or additional meters.
Can a successful extraction still be unusable?
Yes. It may return an error page, a consent page, an irrelevant record or a field without supporting evidence. Validate content and account identity before using it in a CRM note.





