Collect · early access

Not a list. A live stream of the data your business runs on.

Describe your target in plain words. An agent plans the crawl, you approve it, and fresh data from public sources starts flowing in — and keeps flowing. The first 100 sites are free.

Try an example
  • 100 sites free
  • Public sources only
  • Fresh, dated rows
  • No credit card
strategy.yamlexample
goal: Independent dental clinics · DE · online booking
include_if:
- country: DE, language: de
- mentions: Online-Termin | Termin buchen
- booking widget: Doctolib | Dr. Flex | form
- exclude: chains, directories
fields: [name, city, phone, email, booking_tool, fit_score]
Streamlive+23 today
  • +newzahnarzt-mueller.de
  • ~changedsmile-berlin.de
  • +newpraxis-weber.de
420 technologies detected
NamecheapGravity FormsReact RouterTildaCSC Corporate DomainsMobXCookiebotProveSourceGoogle Customer ReviewsAll in One SEOGoogle Cloud DNSRank Math SEOFastmailReddit PixelLottie WebWooCommerce Stripe GatewayGandiYouTube EmbedShop PayTucowsPrivyLeafletWooPaymentsBootstrapAmplitudeGoogle PayAdobe FontsStamped.ioBazaarvoiceJunipAmazon Web ServicesKinstaSectigoWixDynamic YieldImpactAdobe Experience ManagerKustomerLinkedIn Insight TagTrustpilotBeaver BuilderCloudflare RegistrarZoho MailCloudinaryStripeGSAPNew RelicCheckout.comHeapNudgifyWeeblyPlausibleOlarkCustomer.ioDatadog RUMRazorpayConvertKitCloudflare Origin CAOmnisendDriftX PixelXStateBloggerSnap PixelAfterShipWebflowRouteIONOSLogRocketWP RocketMarkMonitorGlobalSignCookieYesAffirmMicrosoft AzureWPBakery Page BuilderUpdraftPlusGoogle Search ConsoleAdaOptinMonsterDiviDrupalWooCommerceComplianzMicrosoft Ads UETTrustPulseAvadaDNSimpleGoogle AdSenseRedirectionFomoProton MailJudge.meNostojsDelivrWebstudioGoogle Identity ServicesGoogle AnalyticsAcuity SchedulingWPMLBolt.newDynadotMatomoTidioWordfence SecurityHelp ScoutCriteoElementorOutbrainAli ReviewsGoogle MapsReferralCandyChatraSEOPressBuypassJimdoLucky OrangeContentfulLovableAmazon CloudFrontZendeskOsanoAzure DNSSoftrRefersionLiteSpeed CacheTailwind CSSContact Form 7ReactMailchimpSiteGroundBunny FontsBunny CDNKeyCDNTaboolaNS1OVH DomainsLoyaltyLionGoogle Tag ManagerPushOwlMagentoYITH WooCommerce WishlistUsercentricsaccessiBeMindbodyCloudflare Web AnalyticsMixoMeta PixelThree.jsAwinMapboxFlatsomeJivoChatCarrdOneTrustHugoCloudflare TurnstilePinterest TagJotformRichpanelKlaviyoGlidenopCommerceAuth0Smile.ioManyChatOpenCartiubendaFancyboxGhostHetznerLiveChatFramerImpervaDigiCertimgixreCAPTCHAWistiaGoogle Ads ConversionEcwidMouseflowLiqPayTanStack QueryHubSpotTrusted ShopsPorkbunWooCommerce BookingsjQuery UISitecoreD3.jsGoogle FontsHostingerAkismetPageFlyPayPalW3 Total CacheAdRollLoop ReturnsOneSignalZodVue.jsLokaliseBluehostNuxtContentsquareUmamiAdyenMarketo EngageVercelBubbleGeneratePressLet's EncryptVWOOptimizelyColorboxTablePressCrispTypeformRxJSLiferayMediavineBrazeNext.jsSquarespace DomainsMicrosoft ClarityJustunoWordPressConvert ExperiencesFreshdeskCalendlyCloudflareRudderStackWisepopsAstroPowerReviewseNomRe:amazeSquarespaceDataDomeMailPoetRadix UIFathom AnalyticsAstraPostscriptYotpoWayForPayWP EngineDigitalOceanTealiumSesamiKlevuDudaNamecheap DNSStoryblokCloudflare DNSOVHcloudTermlyWooCommerce PayPal PaymentsTurbopackVidyardBugsnagHostinger DomainsPolylangDripJoomlaSalesforce Commerce CloudSmash BalloonTYPO3OkendoVimeohCaptchaHotjarKlarnaReplitStatsigFeraSezzleFondyDidomiCoinbase CommerceSentryRebuyhtmxGoogle CloudOpenTablePloneOpenAI Ads PixelServer-side tagging (first-party Google tag)BraintreejQueryRecharge SubscriptionsYandex 360DNNAfterpay / ClearpayAB TastyEzoicWeglotDoofinderTikTok PixelKadenceGoDaddyGorgiasYoast SEOMicrosoft 365LooxTypedreamPrestaShopBase44WPFormsApple PayGemPagesReviews.ioAlpine.jsBigCommerceMuxJetpackNetwork SolutionsMixpanelReact DOMIntercomSanity clientGoogle WorkspaceDurableGoDaddy DNSActiveCampaignLoomViteOracle WebCenter SitesSearchaniseMollieAmazon PayBrevoFriendly Captchatawk.toFeefoUserWaySegmentWP Mail SMTPSwiperFriendbuyFullStoryAlgoliaBoost AI Search & DiscoveryShopifyGatsbyGoogle Trust ServicesBitPaySmartlookAttentiveAmazon Route 53NavidiumAmazon Trust ServicesZeroSSL

From one sentence to a stream

Four steps to set it up. The fifth is the point: the crawl keeps running and your data keeps arriving.

  1. 01

    Describe

    Write what you're looking for, the way you'd tell a colleague.

    Dental clinics in Germany with online booking
  2. 02

    Clarify

    If something's missing, the agent asks 3–5 quick questions. A detailed prompt skips this step.

    Clinic size?
    1–3 dentists4–10AnySkip
  3. 03

    Approve the strategy

    You get a crawl plan: sources, filters, fields. Edit it or approve it as is.

    strategy.yaml
    EditApprove
  4. 04

    Get the data

    The crawler starts. Results appear as they're found — the first ones in minutes.

    +zahnarzt-mueller.de+praxis-weber.de+smile-berlin.de37 matched
  5. the product

    Keep it flowing

    The strategy becomes a stream: new matches and changes arrive every day, straight into your CRM or sheet.

    +23 todaynext pass in 6h

You see the plan before a single page is fetched

The agent turns your prompt into a crawl strategy — a concrete document, not a black box. Nothing runs until you approve it.

  • See exactly what we collect

    Where we look, what a site needs to qualify, which fields we keep.

  • Change anything first

    Tighten a filter, add a field, drop a source — before the crawler starts.

  • Run it again and again

    An approved strategy is what the stream runs on. It keeps finding new matches.

strategy.yamlDental clinics · Germany
Awaiting approval
goal:Independent dental clinics in Germany, online booking
sources:
- search: "Zahnarzt Online-Termin {city}" × 80 cities
- directories: public clinic listings
- expand: links from matching sites
include_if:
- country: DE, language: de
- mentions: Online-Termin | Termin buchen
- booking widget: Doctolib | Dr. Flex | form
exclude_if:
- chain or franchise markers
- marketplace and directory pages
fields:
free:[domain, name, city, phone, email, booking_tool, cms]
llm:[specialisations, dentists, languages, fit_score]paid
volume:~1,200 candidates → first 100 free
stream:daily pass, new + changed

Example strategy. Yours is drafted from your prompt.

What lands in your pack

The first 100 sites are free, with everything we can read without an LLM. Paid adds the analysis — and the stream.

What you getFree · 100 sitesPaid · stream
Sites that pass your strategy's filters
thousands
Name, language, country, title and description
Contacts from public pages: email, phone
Social profiles
Technologies: CMS, e-commerce, analytics, hosting
LLM analysis: what they do, segment, services
Custom fields from your prompt: prices, catalogue…
Fit score: how well each site matches the task
Stream: new matches every day
Change events: new tech, new page, site gone
Stream dashboard and digests
preview
Export
CSV
CSV, JSON, Sheets, API

A few rows of a pack

domaincityemailbooking_toolobservedspecialisationspaidfit_scorepaid
zahnarzt-mueller.deMunich[email protected]Doctolib2026-09-23
praxis-weber.deHamburg[email protected]form2026-09-23
smile-berlin.deBerlin[email protected]Doctolib2026-09-22
dentalpraxis-kiel.deKiel[email protected]Dr. Flex2026-09-22
zahnzentrum-lange.deLeipzig[email protected]Doctolib2026-09-21

A list is a dead end. A stream is a process.

A one-off list starts going stale the day you export it. You can't build a process on it. On a stream you can build a factory: fresh companies come in every day, changes are flagged, and your team works from what's true now.

One-off listexported 3 months ago
  • praxis-hoffmann.declosed
  • zahn-arztpraxis.dechanged site
  • dental-koeln.demoved to Shopify
  • smile-muenster.de
  • zahnarzt-braun.declosed
nothing new since
Stream+23 today
filterenrichscore
  • +zahnarzt-mueller.denew · fit 0.92
  • +praxis-weber.denew · fit 0.81
  • +smile-berlin.dechanged · Doctolib
One-off list / archive
DataCrawly stream
Snapshot, stale from day one
Always fresh, every row dated
Same database everyone else bought
Built for your parameters, unique to you
New companies never show up
New matches arrive daily
Changes go unnoticed
Changes become events: new tech, new page, gone
Re-buy and re-dedupe every quarter
Runs on its own, no duplicates
A file on someone's laptop
Plugs into CRM, Sheets, Slack, API
You can't build a process on it
You can build a factory on it

Lead factory

New clinics that match your ICP land in your CRM every morning, with contacts and a fit score.

Market radar

Know when a competitor's customer switches platforms or a new player enters your niche.

Trigger-based outreach

A site adds online booking, starts hiring or launches a store — your sales team gets a ping the same day.

Watch your data flow in

Crawling the live web takes longer than querying a database. So you don't wait for the end: every site that passes your filters shows up the moment it's checked, and the stream keeps going after the first batch.

Stream active · next pass in 6hexample
  1. 4,812discovered
  2. 1,930checked
  3. 214matched
  4. 180enriched

Last 14 days

  • New matches
  • Changes
Show as table
DayNew matchesChanges
1380
2272
3193
4165
5144
6126
793
8117
9135
10108
11126
12159
13117
14236

Live feed

  • +zahnarzt-mueller.deMunich · Doctolib · fit 0.92
  • ~smile-berlin.deDr. Flex → Doctolib
  • +praxis-weber.deHamburg · form · fit 0.81
  • dental-koeln.deoffline 7 days
  • +zahnzentrum-lange.deLeipzig · Doctolib · fit 0.70
  • First results: minutes
  • 100 sites: under an hour
  • Then: new matches every day

What people stream

Pick one to start from — it goes straight into the field above.

Under the hood

Four parts, each doing one job.

  1. Agent

    Turns your prompt into a strategy: queries, sources, filters, fields.

  2. Crawler

    Finds candidates: search, public directories, expansion along links.

  3. Scraper

    Opens every site and extracts fields without an LLM — the same engine behind our technology profiles.

  4. paid

    LLM layer

    Reads the pages and answers the questions in your strategy.

How we collect
  • Public pages only
  • We respect robots.txt
  • No logins, no paywalls
  • Company contacts only — no personal profiles

Pay for the stream, not for a file

Start free with 100 sites. Keep the strategy running when it proves itself.

  • Free

    $0100 sites
    • Crawl strategy from your prompt
    • 100 sites, fields without LLM
    • CSV export
    Start free
  • Stream

    the product
    TBDper month
    • 1–N strategies running continuously
    • New matches and change events every day
    • LLM analysis and custom fields
    • Dashboard, digests, Google Sheets
    Start a stream
  • Factory / API

    Customvolume
    • Many streams, high limits
    • API and webhooks
    • CRM integrations
    • A one-off large pack if you need it
    Talk to us

Pricing coming soon

Questions

What counts as a public source?

Websites anyone can open without logging in, and public directories. No logins, no paywalls.

How fresh is the data?

Every row carries the moment it was observed. On a stream, sites are re-checked on every pass.

What if the strategy is wrong?

You edit it before anything runs. After the first 100 sites you can refine it and run again.

How long does it take?

First results in minutes, the first 100 sites within the hour, large packs over hours — streamed as they are found.

Why a stream and not just a list?

A list goes stale the day you export it and you can't build a process on it. A stream brings new matches every day and flags what changed.

Can I just get a one-off export?

Yes, export CSV at any time. The value is in the flow, but the file is always yours.

Where does the stream go?

The dashboard, CSV or Google Sheets, your CRM, Slack, or webhooks and the API.

How is this different from BuiltWith?

BuiltWith is an archive: the same database for everyone, as of when it was indexed. Collect builds a list for your parameters from the live web, and keeps it current.

Do you collect personal data?

No. Only public company contacts published on company websites.

What should your stream bring in?

Stop buying lists. Build the factory.

Try an example