nullhex

Scraping 6,234 Coaches: Building a Lead Database from Scratch

20 Mar 2026·7 min read·product

CoachSync needed customers. Specifically, golf coaches and venue operators in the UK and Ireland who might actually pay for the thing. Problem was, no comprehensive database of these people existed. The PGA has membership lists, but those are not public. Golf directories existed but were fragmented, incomplete, and often years out of date. Social media profiles were scattered across a dozen platforms.

So we built the database ourselves. Every golf coach, teaching professional, driving range and golf academy in the UK and Ireland. Scraped from public sources, validated, deduplicated, organised. Final count: 6,234 coaches and venues, 3,954 verified email addresses.

The ethical boundaries turned out to matter more than the technical ones.

// the data landscape

Before writing any scraping code, we spent two days mapping where the data actually lives. Where do golf coaches have a public presence? What information is consistently available? Where do sources overlap?

Four categories emerged. Golf directory websites: Golf Monthly's coaching directory, England Golf's find-a-coach tool, the PGA's public professional search, regional golf unions. These gave us names, club affiliations, sometimes contact details. Venue websites: individual golf club and driving range sites listing their teaching staff. Social media: Instagram, Facebook business pages, LinkedIn profiles. Booking platforms: Fore Business and golf-specific booking tools where coaches list availability.

Each source had problems. Directories were comprehensive but stale - coaches who had moved clubs years ago still listed at their old venue. Venue websites were accurate for current staff but missed freelancers. Social media was current but messy, with inconsistent naming and no standardised contact info. Booking platforms had emails but only for coaches using that specific platform.

No single source gave a complete picture. We needed all of them, cross-referenced to build a composite record for each coach. The scraping was the easy part. Data reconciliation was the real work.

// the scraping approach

One purpose-built scraper per major source. Each one designed to respect rate limits and terms of service. We were not trying to DDoS anyone's golf directory. Just reading publicly available pages at a pace that would not register as unusual traffic.

Directory scrapers were straightforward. Most golf directories are paginated lists with consistent HTML. Parse the list page, extract links to individual profiles, visit each one, pull out the structured data. Name, club, qualification level, contact details where available. Standard patterns.

Venue websites were harder. Every golf club has a different site, built by a different agency, with a different structure. No standard "staff page" template exists in the golf industry. Some clubs have a dedicated coaching page. Others bury their pros in "about us." Some only mention coaches in news articles about recent hires. We used targeted URL patterns (/coaching, /lessons, /professionals) combined with content detection to find the right pages.

Social media scraping was the most nuanced. No private profiles, no authentication required. We only collected from public business pages and professional profiles that coaches had explicitly set up to promote their services. A coach who creates a Facebook business page called "John Smith Golf Coaching" is actively publishing that information. We were just reading it systematically instead of manually.

Every scraper output normalised records in a consistent format. The raw data was a mess - different field names, date formats, ways of expressing qualifications. The normalisation layer cleaned it all up before anything hit the database.

// data structure

The final data model for each coach record:

{
  "id": "coach_uk_4821",
  "name": "James Patterson",
  "qualifications": ["PGA Professional", "TPI Certified"],
  "venues": [
    {
      "name": "Sunningdale Golf Club",
      "type": "private_club",
      "county": "Surrey",
      "country": "England",
      "postcode": "SL5 9RR"
    }
  ],
  "contact": {
    "email": "j.patterson@example.com",
    "email_verified": true,
    "phone": null,
    "website": "https://example.com/coaching"
  },
  "social": {
    "instagram": "@jpgolf",
    "facebook": "jpattersoncoaching",
    "linkedin": null
  },
  "sources": [
    "england_golf_directory",
    "venue_website",
    "instagram_business"
  ],
  "confidence_score": 0.92,
  "last_verified": "2026-03-18",
  "status": "active"
}

The confidence_score was the important bit. A coach appearing in three independent sources with consistent data scored high. One source with partial data scored low. This let us prioritise outreach toward records we trusted and flag the uncertain ones for manual review.

The sources array tracked provenance. Every piece of data traceable to its origin. Not just good practice for data quality - essential for GDPR. If a coach asks where we got their information, we can tell them exactly which public sources it came from.

// validation and deduplication

Raw scraping produced roughly 9,000 records. After deduplication: 6,234. That gap tells you how much overlap exists between sources. Nearly a third were duplicates, sometimes appearing in three or four directories with slightly different spellings or outdated club affiliations.

Name matching was the hard part. "James Patterson" at Sunningdale might be "Jim Patterson" on Instagram and "J. Patterson, PGA" in the England Golf directory. We used fuzzy string matching on names, exact matching on venue associations, and geographic proximity to identify records pointing at the same person.

When two records matched, we merged them by taking the most recent data per field. Directory says the coach is at Sunningdale, venue website says they moved to Wentworth - we trusted the venue site. Instagram has an email the directory lacks - we added it. The merge logic was conservative. When in doubt, we kept both records separate and flagged them for human review rather than risk combining two different people.

Email verification was its own pass. About 30% of raw email addresses were invalid - bounced, deactivated, or mistyped on the source site. Every address went through a verification service checking MX records, mailbox existence, and common typo patterns. The 3,954 that survived were ones we could be confident would reach a real inbox.

Venue matching was another headache. Golf clubs have official names, common names, and abbreviated names. "The Royal and Ancient Golf Club of St Andrews" shows up as "R&A", "St Andrews", or "Old Course St Andrews" depending on who wrote it. We built a normalisation layer mapping all known variations to a canonical venue record with postcode, county, and country.

// the results

The final dataset:

  • 6,234 unique coach and venue records across the UK and Ireland
  • 3,954 verified email addresses (63.4% email coverage)
  • 4,187 venue associations mapped to 1,842 unique venues
  • 2,891 social media profiles linked
  • 0.87 average confidence score across all records

Geographic coverage was solid. England had the densest at 4,102 records, then Scotland at 891, Ireland at 634, Wales at 412, Northern Ireland at 195. The distribution closely mirrors the actual distribution of golf facilities in each region, which gave us confidence the dataset was representative rather than skewed toward any particular source.

Qualification data covered about 70% of records. PGA Professional was the most common, followed by PGA Advanced Professional, TPI Certified, and various federation-specific certifications. Useful for segmentation - a PGA Advanced Professional running an academy has very different needs from a club assistant pro giving occasional lessons.

The dataset was not perfect. No scraped dataset ever is. Some records were already out of date by the time we finished collecting. Some coaches had retired. Some venues had closed. But the verification pass caught the worst of it, and the confidence scoring let us focus on records most likely to represent active, reachable professionals.

// ethical considerations

Scraping at this scale forces you to think about ethics. The technical ability to collect data does not automatically give you the right to use it. We set boundaries before the project started and held to them throughout.

Public data only. Every piece of information came from publicly accessible web pages. No scraping behind login walls, no private databases, no leaked data. If a coach listed their email on their public website, fair game. If it was only in a private membership directory, off limits.

GDPR compliance. UK data protection applies to personal data processing even from public sources. We documented our legal basis (legitimate interest for B2B marketing), maintained provenance records, and built a one-step opt-out that immediately and permanently removed any coach from the dataset. Not buried in a settings page. One step.

Respectful intent. Having someone's email does not entitle you to spam them. When outreach begins, emails will be personalised, relevant, and transparent about where the data came from. Clear one-click unsubscribe. Maximum two emails per person - an introduction and one follow-up. After that, silence means no.

Polite scraping. Honest User-Agent strings, robots.txt respected, delays between requests, immediate backoff on rate-limit headers. We were guests on these sites and we behaved like it.

The ethical constraints cost us data. Sources we chose not to scrape because the terms prohibited it. Email addresses we left alone because the context suggested they were not meant for business contact. The dataset would have been larger without these rules. It would not have been better.

// what the data enabled

With the dataset in hand, CoachSync had a clear picture of its total addressable market. Not a guess. Not an estimate from an industry report. An actual list of real people, at real venues, with verified contact information.

The dataset is ready. Outreach has not started yet, but the groundwork means it can be done properly when the time comes. When you know someone's name, venue, qualification level, and the services they offer, you can write something that reads like a personal note rather than a mass blast. That is the whole point of doing the collection work first.

Even before outreach, the data shaped product decisions. Which regions had the highest coach density. Which venue types were underserved by management software. Trends in qualifications and specialisations. It is not just a contact list. It is a map of the market.

The whole project took about two weeks. Python, BeautifulSoup, Playwright for the JavaScript-heavy sites, SurrealDB for storage and querying. Nothing exotic. The value was not in the tooling. It was in the thoroughness of the approach and the quality of the validation pipeline. These are not rows in a spreadsheet. They are people who might become customers, and the collection process should reflect that.