Data Sourcing, Verification & Accuracy Methodology

Learn where our data comes from, how we verify it, what information we don’t collect, and how often we update our published records. Learn where our data comes from, how we verify it, what information we don’t collect, and how often we update our published records.
DataCaptive Data Sources
01 — Sourcing

How We Source Our Data

DataCaptive builds its database from five distinct classes of source. Each class is governed by its own intake rules, refresh schedule, and rights documentation, and every record carries provenance metadata recording which class it came from, when it was collected, and under which jurisdiction. No record is published on the strength of a single source.

Picture of Public Web Data

Public Web Data

Our crawlers read information that organisations publish about themselves: corporate websites and team pages, regulatory and company-registry filings, press releases, funding and hiring announcements, job postings, conference agendas, and professional directories. Crawling respects robots directives and rate limits, and we collect business information only — never content behind a login, paywall, or private group.
How We Use It
Picture of Licensed Datasets

Licensed Datasets

We license bulk business datasets from established data vendors under written agreements that specify permitted use, redistribution rights, and the lawful basis on which the underlying data was gathered. Every prospective licensor completes a data-provenance review before a contract is signed, and licences are re-assessed at renewal.
How We Use It
Picture of Third-party Data Partners

Third-party Data Partners

Specialist providers supply narrow, high-value datasets — phone intelligence, mailing-address validation, technographic detection, and mailbox status. These vendors are used as instruments rather than as sources of record: their output is treated as evidence for a field, never as the field itself.
How We Use It
Picture of Proprietary Research

Proprietary Research

Our in-house research and tele-verification teams confirm the details machines cannot settle. Analysts call organisations to verify reporting lines, department structure, and whether a number reaches the named person; they also resolve records the automated pipeline has flagged as ambiguous.
How We Use It
Picture of Partner & Opt-in Channels

Partner & Opt-in Channels

Publisher networks, webinar and event registrations, content syndication programmes, and our own sign-up forms contribute records where the individual provided their details directly. Consent language, timestamp, and the collecting party are captured at the point of submission and stored alongside the record.
How We Use It

Records that cannot be corroborated by at least one independent source are held in quarantine rather than published, and quarantined records are excluded from every figure on this page. That is why our published counts are lower than our raw collection volume — we would rather publish a smaller database we can defend than a larger one we cannot.

What We Do Not Collect

Our database is built for business-to-business use. The categories below are out of scope by policy, and our intake filters reject them at ingest regardless of source:
02 — Verification

How We Verify Our Data

Every record passes through seven sequential checks. A failure at any tier sends the record back to quarantine — it is never partially published.

The 7-tier Verification Process

Source validation

Before a record is read, the batch it arrived in is checked for a valid licence or consent basis, a documented collection method, and a jurisdiction we are permitted to process. Batches from an unverified source are rejected at the door and never enter the pipeline.

Syntax & format checks

Every field is parsed against its expected shape: emails to RFC 5322, phone numbers to E.164 with country and area validation, addresses against national postal files, and names and titles against our normalisation dictionaries. Malformed, truncated, and placeholder values are discarded outright.

Deduplication & entity resolution

Probabilistic matching on name, employer, domain, and contact details collapses variants of the same person or company into one golden record with a stable, permanent identifier. Company records are then linked into parent, subsidiary, and branch hierarchies.

Cross-source corroboration

Each published field must agree across at least two independent sources. Where sources conflict, the value is chosen by a weighted score combining source reliability, collection recency, and specificity — and the losing value is kept in the record’s history rather than deleted, so later corrections are auditable.

Live deliverability testing

Each domain is checked for MX records and mail-server health, then the specific mailbox is confirmed through an SMTP handshake that stops before message delivery — no email is ever sent to the contact. Catch-all domains, role accounts, and disposable providers are flagged so they can be excluded from exports.

Human review sampling

Research analysts independently re-verify a statistically significant random sample of every batch by hand, including calling direct dials to confirm they reach the named person. If the sample’s error rate exceeds our threshold, the whole batch returns to tier 4 rather than being corrected record by record.

Compliance & suppression screening

Before release, records are screened against our permanent suppression list, national Do-Not-Call registries, and the processing rules of each contact’s jurisdiction. The same screen runs weekly against the live database, so a record that becomes non-contactable after publication is withdrawn without waiting for its next refresh.

03 — Freshness

45-Day

Refresh Cycle

Every published record is re-validated end to end at least once every 45 days — not sampled, and not limited to records customers have touched. Fast-moving fields such as job title, employer, and phone status are monitored continuously between cycles, because roughly a quarter of business contacts change role or employer each year.

04 — Accuracy

Accuracy by Dataset

Each benchmark is measured across the full published dataset at the close of the most recent audit cycle — not a hand-picked sample — and is restated when the next audit completes.
Dataset Benchmark How it is verified
Contact data
95%
Validated through multi-stage automated verification and analyst review at each refresh, with non-deliverable, role, and catch-all addresses flagged and excluded from the benchmark.
Direct dials
85%
Validated through phone intelligence checks that distinguish direct lines from switchboards, with analyst tele-verification on sampled records to confirm the number reaches the named individual.
Email deliverability
85%
Observed accepted-delivery rate across anonymised, aggregated customer campaign returns, counting hard bounces as failures and excluding content- and reputation-based filtering.

Deliverability depends partly on sender reputation and campaign content, which sit outside our control — the 85% figure reflects hard-bounce-free delivery under standard sending practice.

05 — Volume

Database Size

500M+

Verified contacts

130M+

Phone numbers

75M+

Company profiles

100M+

Direct dials

All counts are published, verified records only: quarantined records, retired records, and records held back by suppression screening are excluded. Figures are rounded down and restated after each quarterly audit — which means a number here can go down as well as up.
06 — Compliance

Compliance & Certifications

We process business contact data under the frameworks below, and our controls are reviewed against them annually. Each entry links to the corresponding control documentation or regulator reference.

Picture of ISO 27001:2022
ISO 27001:2022

Certified

Picture of SOC 2 Type II
SOC 2 Type II

Certified

Picture of GDPR
GDPR

Compliant

Picture of PIPEDA
PIPEDA

Compliant

Picture of DPDPA
DPDPA

Compliant

Picture of CCPA
CCPA

Compliant

07 — FAQ

Frequently Asked Questions

How is the 95% accuracy figure calculated?

The 95%+ benchmark comes from our internal data-quality testing and validation processes, corroborated by validation results our customers report back to us — several of which come in above 95%. It is measured across multiple datasets and samples rather than a single record or campaign, and it reflects ongoing quality processes rather than a one-time historical test. Because the figure is evaluated per attribute, it should be read as a benchmark for the fields being supplied, not as a guarantee that every field in every dataset will always measure exactly 95%. Results can vary by geography, industry, persona, and data type, and internal testing documentation is available to substantiate the claim on request.

What happens to records that fail verification?

A failing record leaves the published database entirely: it is excluded from counts, searches, and exports, and customers cannot reach it. It is re-tested at the next 45-day refresh, and a record that fails three consecutive cycles is retired permanently rather than held indefinitely. Because the pipeline is fail-closed, we never publish a partial record with the failing field stripped out.

Do you send test emails to contacts during verification?

No. Validation stops at the SMTP handshake, which asks the receiving mail server whether the mailbox exists and accepts mail, then closes the connection before any message body is transmitted. The contact receives nothing and sees nothing in their inbox.

How do you handle opt-out and deletion requests?

Deletion, opt-out, and access requests are actioned within 30 days of verification, and in most cases considerably sooner. The record’s stable identifier is written to a permanent suppression list, which every ingest batch and every weekly screen is checked against — so the same person cannot re-enter the database later through a different source. We retain only the minimum needed to keep the suppression effective.

Why is deliverability (85%) lower than email accuracy (95%)?

They measure different things. Accuracy tells you the attribute we supplied is correct — that the address is real and reachable. Deliverability tells you whether a specific message actually landed — which additionally depends on your sending domain’s reputation, SPF/DKIM/DMARC setup, sending volume and cadence, and the content of the email itself. Those factors belong to the sender, so the gap between the two figures is the space that sending practice occupies, not a defect in the data.

Can I test a sample before purchasing?

Yes. We provide a sample file drawn from live production data — not a demo set — matched to the segment you intend to buy, so you can run it through your own validation or enrichment tooling and compare the result against the benchmarks on this page before committing.

Take control of your data

Want to stop receiving communications from DataCaptive? Unsubscribe from our marketing emails anytime using our simple opt-out process.
Exit intend pop up hero
Wait!
Free sample data available
Think no more, first try it and then buy it!
Exit Intend Pop Up

If you don't have a business email, click here






Call DataCaptive