Skip to main content
A competitor raises their prices. Vector checks their pricing page, gets a wall of anti-bot protection instead of the page, and reports what it honestly found: nothing changed. You learn about the price rise a month later, from a customer. Web scrapers prevent that. Monitoring works by downloading a competitor’s page and comparing it with the last copy, so a page it can’t download is a page it can’t watch. Ordinary sites open for any automated fetch, but a good part of the web blocks them — and pricing pages are among the best defended. Scraping services get those pages anyway. They’re worth setting up for three reasons:
  • Coverage. Pages behind anti-bot protection, or built by JavaScript in the browser, only open for a dedicated scraper.
  • Cost. Services are tried cheapest first, and the chain stops at the first success, so an expensive scraper only runs on the sites that need it.
  • Trust. Vector gets the page itself, not a summary of it — which is what makes diffs, archives, and cross-checking between two services possible.
The chain collects pages for scheduled monitoring and pricing scans. Pricing scans try Firecrawl first whenever it has a key. The Help/Docs and Change log crawls in a competitor’s dossier always use Firecrawl, so they need its key.

Choose which services to use

Open Settings and go to the Web scrapers tab. There are four services; the counter on the tab shows how many are ready to use. Out of the box, Jina Reader is tried first, then Firecrawl, then Bright Data; built-in web fetch is off until you enable it.

Add a service to the chain

Each service is a card. Click Enable to put it in the chain, paste its API key, and click Save key. For Bright Data, also enter the name of your Web Unlocker zone in Zone.
The Web scrapers tab in Settings, with a card per service showing its status badge, trust class, arrows for ordering, and an API key field

The chain, cheapest first: each card can be enabled, moved, and given its key

A service that’s enabled but missing a key it needs is marked incomplete and skipped, so a half-finished setup slows nothing down and costs nothing. Disable takes a service out of the chain; Remove key deletes its key.

Put the cheapest first

A fetch tries the enabled services from the top and stops at the first one that returns the page. The arrows on each card move it up or down; an enabled service shows its place, such as enabled #1 for the one tried first. Order is what keeps the cost down. With the cheapest first, an expensive scraper only runs on sites that blocked the cheaper ones — if the first service got the page, nothing below it is called. A service counts as failed, and the fetch moves on, when it’s blocked, errors, times out, runs out of quota, or returns an empty or placeholder page instead of real content. Each service gets 60 seconds, and a page is read up to 200,000 bytes; anything longer is cut off. The tab shows both limits. Whichever service succeeds, Vector turns its output into the same clean format before handing it to your AI provider to analyze.

Check that the chain works

Under Test fetch at the bottom of the tab, paste a URL and run it. It’s a real fetch, so it costs as much as one.
  • retrieved — you see a preview of the content, what the fetch cost, and the escalation path: which services were tried, and why each one before the winner failed.
  • unretrievable — you see the error and the escalation path. The error never exposes a key.

Guard against pages that lie to bots

Some sites don’t block scrapers; they show them plausible but fake content instead. Anti-poisoning protects your monitoring from that: when a change is found by a datacenter service, Vector re-checks it with a human-like one and runs a plausibility check on what came back. It’s on by default. With it off, there’s no cross-checking and nothing waits for your review.
The Anti-poisoning section in Settings, with high-stakes page type chips and a datacenter or human-like trust class per service

Anti-poisoning: which pages are high-stakes, and which services count as human-like

Two settings decide how strict it is:
  • High-stakes page types — where a low-confidence change waits for your confirmation instead of updating anything on its own. Choose from pricing, changelog, releasenotes and blog; only pricing is marked by default.
  • Fetcher trust class — whether each service counts as datacenter or human-like. A change is confirmed by agreement across the two classes. Bright Data is human-like by default; the rest are datacenter.
What the confidence levels mean, and how to clear a change that’s waiting on you, is in Finding confidence and feedback.

What happens when no service gets the page

Vector doesn’t invent content or stop. It marks that source temporarily unretrievable, with the reason, and carries on with the rest. The source is tried again sooner than usual — about an hour later — and goes back to its normal schedule once a fetch succeeds.

See what fetching costs

Every fetch is recorded in your usage log under the service that produced the page, next to AI costs. When Vector works down the chain for one page, the cost includes every attempt that was charged, not only the one that succeeded. Each successful fetch is also kept as a new version of that source, which is how Vector compares it with the previous one and keeps a history for you.

Where to go next

Choose a model per function

Which AI model analyzes the pages.

Put a competitor under monitoring

What these services are collecting for.