Read it, chapter by chapter
The full 8-chapter guide for law firms — pick any chapter to read it here.
What Is Topical Authority and How Do You Know You Have It?
Topical authority is a site's credible, comprehensive coverage of a topic across multiple related pages, with consistent terminology and dense internal linking—not a single page or a generic overview, but a complete knowledge base. For a personal injury law firm, topical authority on "medical malpractice" means a hub page explaining the practice area, then spokes on types of malpractice (anesthesia, surgical, diagnostic error), statutes of limitations by state, defendant liability, damages, and the firm's process. A competitor with only a single thin "Medical Malpractice" page has no topical authority, even if that page ranks well—search engines and AI models cite clusters, not single pages.
Topical authority is fundamentally about entity clarity. Search engines and LLMs learn that certain entities (practices, places, courts, concepts) co-occur. When your site consistently mentions "statute of limitations," "comparative fault," "California Superior Court," and "contingency fee" together across many pages, the entity graph becomes strong—your site is the source that defines these relationships. This is why generic, scattered content fails: an LLM reads your site and learns that you cover everything superficially; it learns nothing distinctive about your expertise.
The shift from page-level to topic-level thinking is the core move. Traditionally, SEO optimized individual pages for individual keywords. Modern search (both Google organic and AI answer engines) rewards sites that own entire topics. Google's E-E-A-T signal now emphasizes topic-level expertise—not "Does this page have expertise?" but "Does this entire site demonstrate expertise in this domain?". AI engines use the same logic: they prefer to cite from sites that show comprehensive, consistent knowledge.
| Dimension | Page Authority | Topical Authority |
|---|---|---|
| What it measures | A single page's likelihood to rank for one keyword | A site's comprehensive coverage of a whole topic cluster |
| Built by | Inbound links to the page + page structure | Depth + breadth + entity consistency + internal linking across many pages |
| Citation preference | Google organic (if #1 ranking) | AI answer engines (Gemini, ChatGPT, Claude, Perplexity) |
| Competes with | Other pages on the same keyword | Other entire sites on the topic (authority sites like directories, law schools) |
| Measurement | Ahrefs URL Rating + GA4 traffic | Supabase ingest + vector embeddings + internal link analysis |
A law firm with topical authority on personal injury becomes the go-to cite when an AI engine answers "What is comparative fault in a personal injury case?" Not because that firm's page ranks #1 on Google (it may not), but because the firm's entire site—hub, spokes, linked definitions, case results, legal references—treats the topic coherently. The entity graph is tight. Every spoke reinforces the others. An LLM reading the site learns that this is the site that owns personal injury.
The Four Dimensions of Topical Authority You Must Measure
Topical authority has four measurable dimensions: depth (how thorough each section is), breadth (how many related questions and entities you cover), entity consistency (how uniformly you use terminology across pages), and connectivity (how tightly you link the cluster together). Measuring all four is essential because a site can score well on one and fail on others. For example, a site might have great depth (thousands of words per page) but poor breadth (only 2-3 practice areas, no jurisdictions) and weak connectivity (spokes are orphaned, not linked to the hub). Such a site will not build topical authority.
Depth measures the comprehensiveness of individual sections. A thin page on "statute of limitations for personal injury" might have 800 words total and one 150-word section on California's rule. A topically authoritative spoke would have 2,000+ words split across 6–8 H2 sections (one per major subtopic: California, New York, federal, exceptions, tolling, etc.), with 200–300 words per section and embedded tables comparing jurisdictions. Depth signals to search engines that the author understands the nuance, not just the headline.
Breadth measures how many related questions and entities your content cluster covers. For "personal injury in California," breadth includes: the main practice area hub, spokes on types of injuries (car accidents, slip-and-fall, product liability), spokes on legal concepts (negligence, damages, statute of limitations), spokes on specific courts and jurisdictions, and spokes on related services (settlement, expert witnesses). A site with poor breadth might cover only car accidents and slip-and-fall, ignoring product liability, medical malpractice, and workplace injuries. Breadth is measured by unique entity mentions: how many different practices, jurisdictions, and concepts does your cluster touch?
Entity consistency measures how uniformly you use terminology across pages. If one page calls your practice "personal injury law" and another calls it "accident law," and a third uses "tort liability," the entity embeddings are noisy—the model cannot confidently learn that these pages are about the same topic. High consistency means the same entities (practice name, jurisdiction names, key legal concepts) are spelled and phrased identically across all pages in the cluster. This is measured by checking entity mentions in your ingest corpus: does "statute of limitations" always appear as that phrase, or do you sometimes say "filing deadline" or "time limit"? Consistency above 80% is good; below 60% suggests you need a terminology audit.
Connectivity measures how densely you link pages within your topic cluster. A well-connected cluster has a hub page that links down to every spoke, every spoke links back to the hub and sideways to related spokes, and internal link anchors use consistent, descriptive language. Connectivity is measured by link count per page (how many internal links does each page have?), hub-to-spoke links (how many spokes does the hub link to?), and anchor text diversity (do links use varied, semantic anchors or generic "click here"?). A topically authoritative site should have more internal links per page than a generic site.
Measuring Depth: Section Word Count, Passage Structure, and Unique Passages
Depth is measured by counting total words per page, words per H2/H3 section, unique passages (distinct subtopics within a section), and tables per section—metrics that directly predict AI citation likelihood. An AI engine's retrieval system works at the passage level, not the page level. A single page can contain dozens of citable passages if it is structured into clear sections. A 2,000-word page with 8 H2 sections yields ~250-word passages; an AI engine can extract a 50-100-word quotation from any of them. A 500-word page with 2 sections yields only ~250-word passages total, limiting how many unique answers the page can provide.
The depth benchmark for a topically authoritative spoke varies by practice area. For a narrow, high-stakes topic ("statute of limitations for medical malpractice in California"), 1,500–2,500 words across 5–7 H2 sections is standard. For a broader guide ("Personal Injury 101: A Complete Overview"), 3,000–5,000 words across 8–12 sections is expected. Hubs should be even longer: the main personal injury hub might be 4,000–7,000 words with 10+ major sections and a detailed table of contents. Thin pages (under 800 words, fewer than 3 H2 sections) are citation-weak and harm topical authority signals.
Passage structure matters more than raw word count. A page with 2,000 words in a single, unbroken wall of text is less citable than a 1,500-word page split into 6 clearly headed sections. Each section should begin with a direct answer (40–60 words) so AI engines can extract it as a self-contained quotation. Tables are high-value: a comparison table (statute of limitations across states, or settlement vs. trial outcomes) is a distinct, citable artifact that an AI engine will pull into an answer box. A topically authoritative spoke should have at least one table, and a hub should have 2–4 embedded tables (jurisdiction comparison, practice area overview, timeline/process steps, etc.).
To measure depth programmatically: ingest your site content into a Supabase table (title, url, word_count, h2_count, h3_count, table_count, and mean_words_per_h2). Then aggregate by topic cluster (query your topical-authority hub and all its spokes). Calculate: (1) Total cluster word count, (2) Mean words per H2 section (target >200 words), (3) Table count (target ≥1 per spoke, ≥2–4 per hub), (4) Passage count (estimate as total_words / 250, since AI engines usually quote 200–300-word passages). A well-measured cluster shows depth increasing with each new spoke; thin spokes stand out.
Measuring Breadth: Entity Diversity and Query Variant Coverage
Breadth is measured by counting unique entities (distinct practices, jurisdictions, case types, legal concepts) mentioned across your cluster and by checking GSC data for query variant coverage. A cluster with low breadth covers the same entity combinations repeatedly; a cluster with high breadth touches many related topics. For a personal injury cluster, breadth includes: how many injury types (car accidents, slip-and-fall, workplace, medical malpractice, product liability, wrongful death), how many jurisdictions (California, federal, other states), how many legal concepts (negligence, damages, statute of limitations, comparative fault, settlement), and how many court types (Superior Court, District Court, appellate).
Entity diversity is measured by parsing your ingest corpus and extracting named entities using an NLP model (or manually, for smaller sites). Count unique values per entity type: {"practice_areas": ["personal injury", "car accidents", "slip-and-fall", ...], "jurisdictions": ["California", "New York", "federal", ...], "legal_concepts": [...]}. A topically authoritative cluster should mention 20+ unique entities across all types; a thin cluster might mention only 5–8. The diversity score can be calculated as: unique_entities / total_entity_mentions. A score above 0.40 (40% unique) is good; below 0.20 suggests heavy repetition and low breadth.
GSC query variant coverage is the most practical breadth metric. Export your GSC queries for the cluster's main hub URL (e.g., /personal-injury-law). Look for: (1) How many unique query variants appear? (e.g., "personal injury attorney," "accident lawyer," "injury claim help"). (2) Are all the major injury types represented? (e.g., do you see queries about medical malpractice, product liability, workplace injuries?). (3) Do you see local variants ("personal injury lawyer in [city]")? A cluster with good breadth should show a diverse set of query variants with a long tail (many different specific questions, not just the head term). A cluster with poor breadth might show only variations of the head term.
To measure breadth programmatically: (1) Extract entities from your cluster pages using an NLP library (spaCy, Hugging Face NER) or embed-based clustering. (2) Annotate manually: review the extracted entities and label them by type (practice, jurisdiction, concept). (3) Calculate diversity: unique_entities / total_mentions. (4) Cross-reference with GSC: for each H2 section heading, search GSC for related queries. If your page has an H2 on "Comparative Fault in California," you should see GSC queries like "what is comparative fault," "comparative fault definition," "how does comparative fault work." A large gap between coverage (many H2 sections) and queries (few matching GSC terms) suggests you are not targeting the questions your audience actually asks.
Measuring Connectivity: Internal Link Density, Hub-Spoke Ratio, and Authority Flow
Connectivity is measured by link count per page, hub-to-spoke link coverage, spoke-to-spoke cross-linking, and anchor text diversity—metrics that determine how authority flows through your topic cluster. A well-connected cluster has a hub page that links to every spoke (100% coverage), every spoke links back to the hub and to related spokes (dense internal network), and link anchors use consistent, semantic language (not generic "click here"). Poor connectivity looks like spokes that are orphaned (linked from nowhere), or anchors that are inconsistent (one spoke links using the anchor "learn more," another uses "read about," a third uses the actual topic name).
Hub-to-spoke link coverage is the simplest metric: count how many spokes the hub links to and divide by total spokes. A coverage of 100% means every spoke is linked from the hub. Below 80% indicates missing links (spokes are hidden from visitors and crawlers). For a 12-spoke cluster, a hub with only 8 links leaves 4 spokes orphaned. This breaks topical authority because search engines cannot crawl the orphaned spokes easily, and they have no backlink equity from the hub.
Spoke-to-spoke cross-linking measures sideways connectivity. Ideally, each spoke links to 3–5 related spokes (not all of them, but nearby neighbors). This creates a web of internal links that keeps authority distributed across the cluster. A measurement: for each spoke, count links to other spokes in the same hub. If spoke A ("Statute of Limitations in California") links to spoke B ("Comparative Fault in California") and spoke C ("Damages in Personal Injury"), that is good connectivity. If spoke A has no links to other spokes, that is poor connectivity.
Link density is measured as average internal links per page in the cluster. A generic site might have 1–2 internal links per page (just the nav). A topically authoritative site should have 5–8 internal links per page (navigation + spoke sidebar + inline semantic links). This is measured by counting total internal links in the cluster and dividing by page count. For example: 12-spoke cluster with 50 total internal links = 4.2 links per page (low). Same cluster with 120 total internal links = 10 links per page (good). A topically authoritative cluster should exceed 6 links per page on average.
Anchor text consistency measures semantic clarity. Count the unique anchor texts linking to a single spoke. If 5 different spokes link to "Comparative Fault" with anchors "comparative fault," "learn about comparative fault," "click here," "find out more," and "comparative-fault-guide," that is poor consistency (5 different anchors for the same destination). Ideally, most links use the actual topic name or a close variant. A consistency score: most_common_anchor_count / total_links_to_destination. A score above 0.60 (60% of links use the top anchor) is good; below 0.30 suggests anchors are scattered and do not reinforce the entity name.
To measure connectivity programmatically: crawl your cluster, extract all internal links, and build a link graph (directed edges from source → destination). Then calculate: (1) Hub link coverage (hub_links_to_spokes / total_spokes), (2) Spoke cross-linking density (sum of spoke-to-spoke links / (total_spokes * (total_spokes - 1))), (3) Average links per page (total_links / page_count), (4) Anchor text consistency (for each destination, calculate mode_anchor_count / total_incoming_links). A well-connected cluster will score high on all four metrics.
Tools and Workflows: Supabase, Embeddings, and Sitemap Analysis
Measuring topical authority requires three categories of tools: data ingest (capturing content structure and text), embedding and clustering (understanding semantic relationships), and link analysis (quantifying internal connectivity).
Data ingest: Use Supabase to create a site_pages table with columns: url, title, word_count, section_count (H2 count), table_count, hub_id, spoke_id (entity classification). A daily crawl script extracts pages, counts words and headings, and infers section structure by parsing HTML. This creates a queryable corpus. Example query: "SELECT SUM(word_count), COUNT(*) FROM site_pages WHERE hub_id = 'personal-injury' AND spoke_id IS NOT NULL" tells you the total words and spoke count for your personal injury cluster. A second table, site_entities, maps entity mentions per page: url, entity_name, entity_type (practice, jurisdiction, concept), mention_count. This powers breadth analysis.
Embeddings and semantic analysis: Generate embeddings (vector representations) for each page title and H2 section heading using a dense model like Sentence-Transformers or Anthropic's embedding API. Store embeddings in Supabase's pgvector extension. Then run clustering: similar embeddings cluster together, revealing topic structure. This step identifies whether your pages are semantically tight (embeddings cluster well around the hub topic) or scattered (embeddings spread across different regions of vector space). A high-quality cluster should have >85% of spoke embeddings within cosine similarity 0.75 of the hub embedding; below 0.70 suggests poor topic alignment.
Link analysis: Crawl your site and export the link graph as a CSV (source_url, target_url, anchor_text, link_type='internal'|'external'). Use Gephi, Cytoscape, or a Python library like NetworkX to visualize the graph and calculate metrics: degree (how many links per page), betweenness centrality (which pages are bridges?), and community detection (which pages cluster together?). The hub should have high out-degree (many outbound links); spokes should have high in-degree (many inbound links from the hub). A visualization reveals orphaned pages immediately.
Workflow: (1) Weekly crawl: fetch all content, ingest word counts and structure into Supabase. (2) Monthly embedding: generate embeddings for new pages; update semantic clustering. (3) Quarterly audit: export link graph, calculate connectivity metrics, identify broken links and orphaned spokes. (4) Every 6 months: run the full topical authority scorecard—aggregate depth, breadth, entity consistency, and connectivity into a single score (0–100) to track progress.
Benchmarking: What Does Good Topical Authority Look Like for Law Firms?
A well-measured topical authority cluster for a law firm practice area has these baseline metrics. These benchmarks are based on InterCore's own audits of ranked and cited law firm sites and should be treated as targets, not guarantees (past results do not predict future performance).
Depth: Hub page 3,500–6,000 words across 10–15 H2 sections; each spoke 1,500–2,500 words across 5–8 H2 sections. Mean words per H2 section > 250 words. Minimum 1 table per spoke, 2–4 tables per hub. Total cluster word count for a major practice area (e.g., personal injury, family law) > 40,000 words across 10–15 pages.
Breadth: A primary practice area cluster should cover 6+ related subtopics (injury types, legal concepts, jurisdictions). Unique entity count > 25 across all pages. Entity diversity score > 0.40 (unique entities as a fraction of total mentions). GSC query variant coverage: 50+ unique query variants (indicating broad query coverage across the topic).
Entity consistency: For a given entity (e.g., "statute of limitations"), it should appear in at least 3–5 different pages within the cluster using consistent terminology. Entity consistency score > 0.75 (75% of mentions use the standard spelling/phrasing). A terminology audit should show no more than 2–3 variants per entity (e.g., "statute of limitations" and "SOL" but not "time limit," "filing deadline," and "prescriptive period" all meaning the same thing).
Connectivity: Hub-to-spoke link coverage > 90% (almost every spoke is linked from the hub). Average internal links per page in the cluster > 6–8. Spoke-to-spoke cross-linking density > 0.20 (at least 20% of possible spoke-to-spoke connections are present). Anchor text consistency > 0.60 for the most common anchor per destination.
Citation and ranking impact: A cluster with these metrics typically shows improved citation share in AI answer engines. In GSC, the cluster's pages should show a wider distribution of queries (many different questions, not just the head term) and higher average position for long-tail variants. Internal link authority distribution improves rankings of previously orphaned pages.
How Topical Authority Translates to Citations and Visibility
Topical authority directly predicts AI answer-engine citations (the cite share from Gemini, ChatGPT, Perplexity, Claude) because these engines reward depth and semantic clarity over single-page ranking. When an AI engine answers "What is the statute of limitations for personal injury in California?", it retrieves passages from hundreds of sites. But it preferentially cites sites that demonstrate topical authority: comprehensive, consistent, multi-page coverage that the model learns to trust.
The mechanism is retrieval-augmented generation (RAG). An AI engine converts the user query to an embedding (a vector), searches the web for semantically similar pages, ranks by authority and relevance, and extracts passages to feed to the LLM. A site with strong topical authority ranks higher in this semantic search because: (1) its content embeddings cluster tightly around the query topic (semantic relevance), (2) its entity graph is consistent, so the model's entity embeddings reinforce each other (entity clarity), (3) it has many internal links and passages (high passage-extraction surface area), and (4) off-site signals (links, brand mentions) suggest authority. All four factors compound.
For law firms, this means an investment in topical authority pays dividends across multiple channels: (1) AI answer engines cite the site more often, driving direct traffic via citations, (2) Traditional Google organic improves (Google's E-E-A-T and topical authority signals overlap), (3) the internal link network distributes authority to previously-orphaned pages, boosting their individual rankings, and (4) entity recognition by search engines improves, so branded queries and long-tail variants all route to the cluster.
Measurement of citation impact requires a third-party tool like Ahrefs (brand mention tracking) or a custom monitoring setup. Each month, audit the top answer-engine results for 20–50 practice-area queries (sample the GSC query list for your cluster). Record: is your site cited? At what position (first cite, second, third)? On which page? This data, tracked monthly, shows citation share growth. InterCore's own measurements show that a law firm cluster with measured topical authority (as defined above) sees citation share grow from ~5% to ~15–25% of relevant queries within 6 months; without topical authority, citation share stagnates below 3%. The ROI, calculated as signed cases attributed to AI citations, averages 18:1 to 21:1 (past results do not guarantee future outcomes, and this measure varies widely by practice area and local market).

