How to Scrape Competitor Landing Pages for Semantic Keyword Patterns
How to Scrape Competitor Landing Pages for Semantic Keyword Patterns Introduction Competitor landing pages contain your most valuable keyword research data. But manually reviewing competitor content misses the patterns that matter. Semantic keyword extraction — analyzing the relationships between keywords, themes, and topics — reveals how competitors structure their authority. By scraping competitor pages at scale, you can identify the exact keyword families, topic clusters, and content gaps that drive their rankings. What Semantic Keyword Patterns Are and Why They Matter Semantic keyword patterns go beyond simple keyword frequency. They capture the relationships between keywords, the themes that connect them, and the context in which terms appear. A single landing page might use “real estate attorney,” “property lawyer,” and “closing counsel” interchangeably. These are not separate keywords. They are semantic variants of the same underlying topic. When you scrape competitor landing pages for semantic patterns, you are not just collecting keyword lists. You are building a map of how competitors organize their topical authority. This map reveals which themes they prioritize, which concepts they treat as related, and which specific phrasing they use to match search intent. The core difference between traditional keyword extraction and semantic pattern analysis is grouping. Traditional extraction gives you a flat list. Semantic analysis groups variants into themes, identifies which themes appear across multiple competitors, and surfaces the concepts that define your competitive landscape. Scraping Competitor Landing Pages: What to Extract Before analyzing semantic patterns, you need structured data from competitor pages. The essential fields for semantic analysis include the full page title, all heading elements from H1 through H3, the meta description, visible body text excluding navigation and footer content, and any structured data or schema markup present on the page. For multi-market analysis across the USA, Germany, United Kingdom, France, Italy, Russia, Spain, Netherlands, Switzerland, Poland, Ireland, Australia, Canada, Thailand, and Hong Kong, run separate scrapes for each target location. Semantic patterns vary by language, cultural context, and local search behavior. A keyword theme that appears consistently in US competitor pages may be entirely absent from German competitors. The technical approach can range from custom scripts using Python libraries like BeautifulSoup or Scrapy to managed scraping workflows using platforms like Decodo or the CustomJS Scraper node in n8n, which fetch raw HTML and extract key SEO elements including title, headings, and meta data. Extracting Keywords and N-Grams from Scraped Content Once you have the raw content, the next step is extracting keyword phrases at multiple lengths. Unigrams — single words — are too noisy for semantic analysis. Focus on n-grams, which are phrases of two to four words. Bigrams like “real estate” and trigrams like “real estate attorney” capture the specific language competitors use. The Apify SEO Keyword Extractor uses a transformer-based model to extract multi-word keyphrases from page content, filters out numeric strings and technical junk, and keeps the most relevant two to four word keyphrases per page. The Apify Analyze Website Content tool extracts the most frequent n-grams across two to four words and identifies keywords from HTML metadata. For local or practice-area SEO, pay close attention to geo plus service combinations. Phrases like “fort lauderdale real estate lawyer” or “west palm beach probate attorney” reveal the specific location-modifier patterns competitors target. These combinations are often invisible to traditional keyword tools but appear clearly in scraped competitor content. Clustering Keywords into Semantic Families The most valuable output from semantic analysis is keyword families — groups of related phrases that represent the same underlying concept. Clustering similar phrases across multiple competitor pages reveals which concepts dominate your market. The process involves identifying all extracted phrases, calculating similarity between phrases using token-set matching or Levenshtein distance, grouping phrases that share core tokens, and for each group, selecting a representative phrase. A group containing “florida real estate attorney,” “florida real estate lawyers,” and “florida real estate law” would cluster under “florida real estate attorney” as the representative. Tools like the SEO Keyword Extractor compute cross-site keyword families by clustering similar phrases across multiple domains. The output includes the group representative, all variant keywords in the group, the number of distinct keywords in the group, and which competitor sites use each variant. This tells you not just what competitors are targeting, but how consistently they target it. Identifying Common Cross-Site Themes Phrases that appear across multiple competitor sites are signals of market standards. If three or four competitors all target variations of “real estate attorney near me,” that concept is not optional for your content strategy. The SEO Keyword Extractor calculates n-gram statistics for phrases that appear on at least three different sites, treating these as strong cross-site themes. For each n-gram, the tool returns the phrase text, the number of sites using it, the total count across pages, and sample keywords showing the full phrase variants. For example, analyzing competitor sites in the legal industry might reveal that the trigram “fort lauderdale real” appears across four competitor sites with sample keywords including “fort lauderdale real estate,” “lauderdale real estate lawyer,” and “lauderdale real estate attorneys”. This tells you that the combination of location and practice area is a mandatory theme in your market. Building Ranked Keyword Themes The final stage of semantic analysis is merging similar keyword families into higher-level themes and ranking them by importance. A keyword theme represents a complete topic area that your content should address. The SEO Keyword Extractor builds themes by constructing a graph of keyword groups connected by high Jaccard similarity — meaning groups that share a high proportion of their word sets — then collapsing connected components into themes. Each theme includes a primary keyword representing the best phrase for the theme, a score indicating theme strength based on cross-site importance and cohesion, the number of distinct keyword variants in the theme, and the complete list of all variant phrases. A theme with primary keyword “florida real estate attorney,” a score of 0.95, three sites in the theme, and variants including “florida real estate law” and