TL;DR
- Lead generation is a sourcing and interpretation problem, not a writing problem. The leverage is in reading unstructured evidence and turning it into a structured record, which is exactly what a language model is good at.
- Do not use it as a database. Contact data must come from a provider or a live source. A model asked for contact details will produce plausible ones, and plausible contact data is worse than none.
- Interpretation is the real win. Reading a job posting, a review, or a project brief and extracting what it implies about a company's problem is work that previously required a person.
- Qualification only works against a written, observable profile. If a criterion cannot be verified from public information in about two minutes, it cannot be scored.
- Signals decay at very different rates, from 24 to 72 hours for launches to three to nine months for funding, and treating them as equivalent is the most common structural error.
Contents
- What a language model is and is not good for here
- Sourcing: turning unstructured evidence into candidates
- Enrichment: interpretation rather than lookup
- Committee mapping
- Qualification against a written profile
- Handling signal decay
- Guardrails and data handling
- FAQ: Using Claude for Lead Generation
What a language model is and is not good for here
The single most useful distinction in this whole subject: a language model is an interpreter, not a database.
Good at: reading unstructured text and extracting structure from it. A job posting becomes a set of implied problems and an implied timeline. A negative review becomes a switching signal with a named incumbent. A conference exhibitor list becomes a segmented account list. A project brief becomes a scoped need with a budget range.
Bad at: recalling facts about specific companies and people. Ask for a contact's email address and you will receive a plausible one, formatted correctly, that may not exist. Ask when a company raised its last round and you may receive a confident wrong date.
The boundary is clean and it should be enforced structurally rather than by discipline: contact and firmographic data comes from a provider or a live source, interpretation comes from the model, and every claim carries the source it came from.
| Task | Suitable | Why |
|---|---|---|
| Reading a job posting for implied need | Yes | Interpretation |
| Classifying a review as a switching signal | Yes | Interpretation |
| Structuring an exhibitor list | Yes | Extraction |
| Recalling an email address | No | Fabrication risk |
| Recalling headcount or funding | No | Fabrication risk |
| Verifying a claim you already have | Partly | Needs a live source |
Teams that get poor results almost always crossed this line somewhere, usually without noticing, and then concluded the approach does not work.
Sourcing: turning unstructured evidence into candidates
Firmographic filtering is not a competitive advantage, because everyone with a database subscription can build the same list. What differentiates a list is the signal attached to each account, and signals live in unstructured text that databases do not index.
The workflow: extract the raw pages from a source, then interpret each one into a structured candidate record.
Job postings. The richest single source. A posting reveals the function being built, the tools mentioned, the seniority of the hire, the problem implied by the responsibilities, and often the budget implied by the salary band. Interpreting a posting into implied need is genuinely hard for rules and easy for a model.
Reviews in your category. A three-star or lower review posted in the last 90 days names the incumbent and the specific dissatisfaction. That is a switching signal with the objection already written out for you.
Project briefs and requests for proposals. A scoped need with a timeline and often a budget. Among the strongest signals available, with a short useful window of roughly 7 to 21 days.
Launch and announcement feeds. Strong but extremely perishable, with a useful window of 24 to 72 hours.
Community and forum discussion. Someone describing the problem you solve in their own words, which is both a signal and the best possible source of message language.
Conference exhibitor and speaker lists. A pre-segmented list of companies that spent money to be seen by a particular audience.
The output of this stage is a candidate record per account with the trigger, the exact source URL, and the trigger date. Nothing else yet.
Enrichment: interpretation rather than lookup
This is the stage that is widely misunderstood. Enrichment in the conventional sense means appending fields from providers, and that is a data problem solved with waterfall logic across multiple providers rather than with a model.
What a model adds is the interpretation layer on top of enriched data.
Implied problem. Given the trigger and the firmographics, what is this company most likely struggling with right now, stated in one sentence.
Implied urgency. Whether the evidence suggests an active project, a forming intention, or nothing in particular. Most accounts fall into the third category and saying so is the useful part.
Incumbent and switching context. What they appear to use now and whether there is evidence of dissatisfaction.
Evidence to reference. Two or three specific, verifiable observations that could appear in a message, each with its source.
Disqualifiers. Anything found that argues against pursuing this account. The most valuable output and the one most often omitted, because a system that never says no scores everything as a fit.
| Layer | Handled by | Output |
|---|---|---|
| Contact and firmographic fields | Providers, waterfall logic | Verified data |
| Trigger detection | Extraction plus interpretation | Signal with source |
| Implied problem and urgency | Interpretation | One-sentence read |
| Evidence selection | Interpretation | Sourced observations |
| Disqualification | Interpretation | Reason to skip |
Committee mapping
Purchases in enterprise and upper mid-market accounts involve several people, and single-threading is among the most common reasons well-sourced pipeline stalls.
What a model does well here: given a company, a trigger, and what you sell, reason about which three to five roles would be involved and why each one cares. That reasoning is genuinely useful and does not require recalling any specific fact.
What it must not do: name the individuals. Names and titles come from a provider or a live source. Use the model to define the roles you need, then fill them from data you can verify.
The output per account should be three to five roles, each with a one-line statement of why this trigger matters to that role specifically. That statement is what makes multi-threaded messaging different from sending the same message to more people, which is the usual failure.
Qualification against a written profile
Qualification works when the criteria are observable and fails when they are not.
Write the profile as observable attributes. If a criterion cannot be verified from public information in about two minutes, it cannot be scored, and including it produces a number that looks precise and means nothing.
Score in bands. Strong fit, possible fit, no fit. Decimal scores imply a resolution the underlying data does not support and invite people to trust the number more than it deserves.
Record the reason with the score. Two or three driving factors alongside every band. A score without a reason cannot be audited and will drift silently.
Calibrate on known outcomes. Run the qualification over accounts you already won and already lost. If closed-won accounts do not score highly, the profile is wrong, and tooling does not fix a wrong profile.
Re-audit quarterly against a human judgement on a sample. Divergence means either the criteria drifted or your market did, and both are worth knowing.
Handling signal decay
The most common structural error in signal-based lead generation is treating all signals as equally durable. They are not, and acting on a stale signal is worse than not acting, because it advertises delayed automation.
| Signal | Predictive strength | Useful window |
|---|---|---|
| Product launch | Strong | 24 to 72 hours |
| Active project brief | Strong | 7 to 21 days |
| Negative review in category | Strong | 30 to 90 days |
| Hiring into the function | Strong | 30 to 90 days |
| Technology stack change | Strong | 30 to 90 days |
| Conference exhibitor | Moderate | Around the event |
| Executive appointment | Moderate | 60 to 180 days |
| Funding round | Moderate qualifier | 3 to 9 months |
| Third-party intent topic surge | Weak alone | Short and noisy |
| Firmographic match | Weak alone | Indefinite |
Two implications worth building into the system.
Store the trigger date and enforce the window automatically. A record whose trigger has expired leaves the sequence rather than being flagged, because flagged items get sent under time pressure.
Match cadence to decay. A 24 to 72 hour signal needs a same-day path from detection to send, which means the pipeline has to run continuously rather than weekly. If you cannot operate at that speed, do not source that signal type, since a late launch reference is actively damaging.
Guardrails and data handling
Never let a model supply contact data. Enforce this in the schema by making contact fields provider-only, not by relying on a guideline.
Every claim carries a source URL. Any claim without one is stripped before a message is drafted, not flagged for review.
Respect source terms and applicable data protection law. Whether a given source may be extracted depends on that platform's terms, the jurisdiction, and how the resulting data is used. Personal data has additional obligations in many jurisdictions, and this is a question for your own legal advice rather than something to infer from an article.
Keep deliverability discipline separate and non-negotiable. Authenticate with SPF, DKIM, and DMARC, include one-click unsubscribe per RFC 8058 and honour it within two days, keep bounce rate under 3%, and keep complaint rate below 0.10% and well away from 0.30%. Bulk sender rules apply from 5,000 messages a day to Gmail.
Do not measure on open rate. Apple Mail Privacy Protection makes it meaningless. Measure reply rate, positive reply rate, and meetings held, against benchmarks of 3.4 to 5.8% average cold email reply with top performers at 10 to 18%, and a positive reply rate around 1 to 2%.
A person reviews every message that reaches a prospect. The failure mode is confident specificity that happens to be wrong, and a fabricated detail reads exactly like a real one.
Clay Waterfall Enrichment Guide: Download Free
Get the practical framework for applying this article to your GTM system.
Download the Clay Waterfall Guide →
FAQ: Using Claude for Lead Generation
Can Claude find leads for me?
Not in the sense of supplying contact records, and this is the distinction that determines whether the approach works. It is an interpreter rather than a database, so asking for an email address or a headcount figure produces something plausible and formatted correctly that may not exist. What it does extremely well is read unstructured evidence such as job postings, reviews, and project briefs and turn it into structured candidate records. Contact and firmographic data must come from a provider or a live source.
What is the highest-value use in lead generation?
Interpretation of unstructured signals. A job posting reveals the function being built, the tools in use, the seniority of the hire, and the problem implied by the responsibilities, which is difficult to capture with rules and straightforward for a model. Similarly, a three-star or lower review posted in the last 90 days is a switching signal that names the incumbent and states the objection for you. This work previously required a person and is the reason signal-based lists outperform firmographic ones.
How should qualification scoring be set up?
Against a profile written entirely in observable attributes, since any criterion that cannot be verified from public information in about two minutes produces precision without meaning. Score in bands rather than decimals, record the two or three driving factors alongside each score so it can be audited, and calibrate by running the scoring over accounts you already won and lost. If your closed-won cohort does not score highly, the profile is wrong rather than the tool. Re-audit quarterly against human judgement on a sample.
How do I handle signals that go stale?
Store the trigger date and enforce the useful window automatically, removing expired records from sequences rather than flagging them, since flagged items get sent under time pressure. Windows vary enormously: product launches last 24 to 72 hours, project briefs 7 to 21 days, hiring and stack changes 30 to 90 days, executive appointments 60 to 180 days, and funding three to nine months as a qualifier only. Match your operating cadence to the fastest signal you source, or stop sourcing it.
Should it map the buying committee?
It should define the roles, not name the individuals. Given a company, a trigger, and what you sell, reasoning about which three to five roles are involved and why each one cares about this specific trigger is useful work requiring no factual recall. The names and titles filling those roles must come from a provider or a live source. The per-role reason is what separates genuine multi-threading from sending the same message to more people, which is the usual failure.
What guardrails are non-negotiable?
Contact fields are provider-only and enforced in the schema rather than by guideline. Every claim carries a source URL, and claims without one are stripped before drafting rather than flagged. A person reviews every message that reaches a prospect. Deliverability discipline stays separate and strict: SPF, DKIM, and DMARC authentication, one-click unsubscribe per RFC 8058 honoured within two days, bounce under 3%, and complaint rate below 0.10%. Source terms and data protection obligations are a matter for your own legal advice.





