Cascade pipeline
Discovery, company-website reading, people enrichment, email verification and CRM suppression, each stage able to be switched off without breaking the ones after it.
Data product · 2026
Design and build
A lead pipeline for a client that finds candidates across four sources, confirms every field against at least two of them, and refuses to hand over anyone already sitting in the client's CRM.
sources cross-referenced per candidate
CRM contact records checked against every run
sources must agree before a field is confirmed
records written to the CRM without human approval
The brief
The client's problem was not finding leads, it was trusting them. Bought contact data arrives confident and partly wrong, and a sales team that gets burned twice stops using the list. On top of that, contacting someone who is already a client, or already said no, costs more than never contacting them at all.
The engine discovers candidates by position, company and location, then runs each one through a cascade where every source that follows both verifies what is already there and fills what is missing. A field counts as confirmed only when at least two sources agree.
It ends by cross-referencing the client's own CRM, which holds around 57,000 contact records, and either drops the match or shows it as a field-level difference for a human to approve.
Deliverables
Discovery, company-website reading, people enrichment, email verification and CRM suppression, each stage able to be switched off without breaking the ones after it.
Every field on every candidate carries the trail of which sources produced it, and a confirmed badge counting only the sources that actually ran.
A weekly full mirror of the client's CRM, an indexed match on email then name plus company then name alone, and a decision on each match that a human signs off before anything is written.
A single-page dashboard for searching, inspecting the source trail behind any candidate and exporting, with searches dispatched as scheduled jobs.
A scoring pass that replays the website-reading stage over past runs and reports how often a delivered lead is actually named on their own employer's site, using no paid calls.
Actions
Problems and solutions
Every project has these. They are more informative than the finished result, so they are on the page.
Problem
A month's budget of paid enrichment credits was emptied in a single day.
Solution
Two guards, both asserted in the test suite so they cannot quietly regress. The paid source was restricted to enriching people already found rather than discovering new ones, which means no names in equals no call out and no spend, and a per-company cap was set to match the one lead per company actually delivered. Letting it discover again is worth around 38 percent more leads at several times the burn, which is a business decision to be taken deliberately, not a default to drift into.
Problem
The CRM offers no search endpoint and no changed-since endpoint, so answering whether a person already exists meant paging the entire database on every run.
Solution
Moved the full pull to a scheduled weekly job that builds and caches an index, so weekday runs cross-reference from memory instead of the network. An empty index is treated as a hard error rather than as a clean result, because the failure mode of a silently empty suppression list is contacting every existing client at once.
Problem
Real contact data was committed to git history once, putting a small number of real email addresses and phone numbers into a permanent record.
Solution
The data directories are gitignored, the database is the store of record rather than any file, and no scheduled job commits anything except a push receipt carrying counts and a timestamp and no personal data at all. Which individuals were delivered is answerable inside the client's own CRM, which has a retention policy, rather than in git history, which does not. The historical exposure is documented in the open rather than quietly cleaned, because removing it from history is a separate decision that has not been made.
Problem
The database shares an instance with an unrelated live product, so a restore to undo a mistake here would roll that product back too.
Solution
Migrations are forward-only and additive with no drops and no renames, and the runner refuses any migration file that names the shared schema, with that refusal asserted in the test suite. The safety comes from the migration being unable to reach the other product, not from remembering to be careful.
Problem
The dashboard had reset and seed buttons sitting one mis-click away from wiping four tables of the client's live data.
Solution
Removed from the interface and replaced with command-line operations that each require an explicit target. There is no wipe-everything path left to click.
Outcome
A lead list the client's sales team can act on without checking it first, where every field carries its own evidence, nobody already in the CRM gets contacted twice, and nothing is written back without a human approving the specific rows.
What it took