AI search is reshaping how B2B SaaS marketing leaders find and shortlist demand generation agencies — the agencies that establish citation visibility now lock in a structural advantage before the rest of the category catches up.
Before the audit measures citation visibility in the B2B demand generation agency space, these three signals tell us whether AI crawlers can access, parse, and trust refinelabs.com at all.
AI search is changing how B2B SaaS marketing leaders discover and evaluate demand generation agencies — the same shift in buyer behavior that Refine Labs built its brand teaching the market about is now reshaping how its own buyers find agencies. Early citations compound: as AI platforms learn which domains to trust for a category, the agencies cited first become the default answer for every buyer who follows. Refine Labs enters this window with a strong brand in a category where no player has yet locked in AI-search dominance — which makes this a timing opportunity, not a recovery project.
This Foundation Review presents the three inputs the audit runs on: the competitive landscape that determines which head-to-head matchups get tested, the buyer personas that determine how queries are phrased and what intent they carry, and the Layer 1 technical baseline that determines whether AI platforms can access and extract your content at all. We're validating these together before the audit runs — every correction you make here changes what the query set measures.
The validation call is a decision-making session, not a readout. Two kinds of decisions get made: input validation — are the right competitors in the right tiers, the right buyers in the right roles, the right capabilities rated honestly? — and engineering triage — which technical fixes start now, before results come back. The Pre-Call Checklist at the end of this document collects every question and task in one place.
What this is This document is the research foundation for a GEO visibility audit of Refine Labs in the B2B demand generation agency category. Everything below — competitors, buyer personas, capability ratings, pain points — was built from your website, review platforms, and third-party agency listings. It drives the buyer queries we'll run across AI platforms, so its accuracy directly determines the audit's accuracy.
What we need from you Read each section and react. Where we got it right, say so. Where we got it wrong — a competitor that never shows up in deals, a persona who doesn't exist in your buying committee, a capability rated too generously or too harshly — correct us. The purple question boxes throughout mark the specific spots where your judgment matters most.
Confidence badges Every item carries a confidence badge. High means multiple corroborating sources or direct evidence from your site. Medium means a single source or an inference — these deserve your closest scrutiny, and they're where the purple questions concentrate.
Validate AI training data still strongly associates Refine Labs with founder Chris Walker, who exited in July 2025 — answers may attribute your positioning to him rather than CEO Megan Bowen. Should the audit test founder-associated queries (e.g., "Chris Walker's agency") as brand variants, or has that association become something to measure as a separate risk? If founder queries still drive discovery, we add them to the brand query cluster; if not, we test whether AI answers have caught up to the leadership change.
5 personas: 2 decision-makers, 1 evaluator, 2 influencers — each drives a distinct cluster of buyer queries in the audit.
Critical review area Personas are the highest-leverage input in this document. Each one generates a distinct set of buyer queries phrased in that person's language and intent. A persona who doesn't actually sit in your buying committee wastes query budget; a missing persona means a whole class of real buyer questions goes unmeasured.
Data sourcing note Role, department, seniority, influence level, veto power, and technical level come directly from the knowledge graph (sourced from your site's case studies and testimonials, review mining, and — where flagged — inference). Role descriptions, buying jobs, and query focus areas are synthesized from those fields to show how each persona will be represented in queries. Correct anything that reads wrong.
→ Does the CMO personally build the agency shortlist, or delegate it to the VP of Demand Gen? If she delegates, we shift weight from C-suite proof queries to mid-funnel comparison queries where the shortlist actually forms.
→ Does the VP of Demand Gen control the agency budget line at any of your target accounts? If yes, we reclassify him as a decision-maker and add budget-justification queries to his cluster.
→ At your smaller accounts, is the Director of Marketing actually the one running the evaluation rather than a VP? If so, we promote her to evaluator and weight her hands-on comparison queries accordingly.
→ This persona is inferred, not observed — does RevOps actually sit in your agency evaluations, or only vet attribution after the contract is signed? If they're not in the deal, we cut this persona and reallocate his attribution queries to the CMO's cluster.
→ At your $50M+ ARR target accounts, does the CEO personally research and approve agency selection, or is the CFO the real budget gatekeeper? If it's the CFO, we replace CEO-style queries with finance-oriented ROI and cost-justification queries — which read very differently in AI search.
Missing personas? These roles sometimes appear in B2B demand generation agency deals — do they show up in yours? CFO / VP Finance (if a $240k+/year retainer triggers formal finance approval, that's a distinct query cluster about agency ROI and cost justification). CRO / VP Sales (your own pain point data shows sales-marketing misalignment — if sales leadership weighs in on agency selection, their skepticism drives different queries). Head of Growth (at product-led SaaS companies, growth sometimes owns paid media instead of marketing). Who else shows up in your deals?
5 primary + 4 secondary competitors identified — tier assignments determine which head-to-head matchups the audit tests.
Why tiers matter Primary competitors each get 6–8 dedicated head-to-head queries — roughly 30–40 queries like "Refine Labs vs Directive for B2B SaaS" or "best B2B demand generation agency for pipeline, not MQLs" — while secondary competitors appear only in category-level queries. We're less certain about Fullfunnel.io's primary tier: it was sourced from third-party alternative listings rather than direct comparison pages (Refine Labs publishes none), and if it rarely appears in your actual deals, moving it to secondary shifts 6–8 queries back into the head-to-head set for a competitor who does.
Validate Three tier questions, in order of consequence: (1) Do prospects actually weigh Metadata.io — software — against hiring you, or is the real alternative always another agency or building in-house? If software never appears in your deals, those queries move to agency head-to-heads. (2) All nine tiers were sourced from third-party listings because Refine Labs publishes no comparison pages — should Fullfunnel.io (primary) and Goldenhour (secondary) swap, based on who you actually see in late-stage evaluations? (3) Is anyone here irrelevant — for instance, does Kalungi's seed-to-Series-B focus ever intersect with your Series-B-and-beyond pipeline, or should it drop entirely? And which agencies that beat you in deals are missing from this list?
12 buyer-level capabilities mapped: 6 strong, 2 moderate, 2 weak, 2 absent — each becomes a capability query phrased in buyer language.
Move from lead generation to demand generation — create demand with buyers who aren't in-market yet instead of just capturing existing search demand
An agency that can actually run LinkedIn and Meta ads for B2B and turn ad spend into qualified pipeline, not just impressions
Google Ads management that captures high-intent demand efficiently and stops wasting budget on junk keywords
High-volume, high-quality ad creative — copy, video, and motion graphics — that stands out in B2B feeds and doesn't fatigue after two weeks
Reporting that ties marketing spend to qualified pipeline, CAC, and win rates — numbers I can put in front of the board
Help sharpening our category positioning and messaging so buyers actually understand and remember what we do
Marketing programs that drive expansion revenue from existing accounts and target named accounts, not just new logo acquisition
An agency that can grow organic search and AI-search visibility alongside paid, so we're not renting all our traffic
Executive thought leadership and organic social content that creates demand on LinkedIn without paying for every impression
Cold outbound, email sequencing, and GTM engineering to complement inbound demand programs
Landing page testing and website conversion optimization so the traffic we buy actually turns into demo requests
Upskill my in-house marketing team on modern demand gen instead of staying dependent on an agency forever
Prioritization The audit tests all 12 capabilities, but competitive differentiation queries will emphasize 3. Six capabilities are rated Strong:
• Demand Creation Strategy
• Paid Social Advertising Management
• Ad Creative & Video Production
• Pipeline Measurement & Revenue Attribution
• Brand Strategy & Positioning
• Marketing Team Education & Enablement
Which of these best represents where Refine Labs wins deals?
Validate Three things: (1) We rated Paid Search Advertising Management "moderate" because multiple third-party guides describe you as "paid social and strategy only," while your pricing page and the myCOI case study say Google Ads is in scope — if paid search is now a core strength, head-to-head queries against KlientBoost and Directive shift from a vulnerability gap to a battleground we measure aggressively. (2) SEO & Organic Search and Outbound & Sales Development are rated absent on multiple corroborating sources — confirm these are genuinely out of scope, since they're the clearest gaps competitors like Powered by Search and NoGood will win queries on. (3) Do Brand Strategy & Positioning and Demand Creation Strategy read as one capability to your buyers, or two distinct purchase conversations? If one, we merge their query clusters.
9 pain points: 4 high, 5 medium severity — the buyer language below is how audit queries will actually be phrased.
Validate Is "Expansion revenue under-marketed" really medium severity for your buyers, or is it a high-stakes conversation at renewal-driven SaaS accounts — and does the buyer language above sound like your actual prospects, or like marketing copy? Also, three pains we see in demand gen agency deals that aren't in this set — do they show up in yours? Agency churn fatigue ("we've burned through three agencies in four years" — skepticism shapes how buyers phrase evaluation queries), in-housing pressure (a CFO pushing to bring paid media in-house rather than renew a $20k+/mo retainer), and slow time-to-impact (demand creation takes quarters to show pipeline, which is hard to defend mid-contract). What's missing?
8 findings from the technical analysis of refinelabs.com: 2 high severity, 4 medium, 2 low — none block crawler access, but the two high-severity items directly suppress your strongest proof content.
Actionable now No critical blockers — AI crawlers can reach the site — but two high-severity findings need attention before the audit measures visibility. Engineering should fix the case-study page template (all 11 success stories ship with the bare title "Refine Labs," no meta descriptions, no OG tags — a single CMS template change) and configure the sitemap generator to emit lastmod timestamps. Content should start the refresh cycle on the ~13 stale high-value pages, beginning with The Attribution Mirage and the Paid Media Benchmarks post that still carries 2024 data. All three can start today without waiting for the validation call.
What we found: Refine Labs' flagship thought-leadership pieces on attribution and measurement — The Attribution Mirage (last modified 2025-06-11), Proving ROI (2025-05-01), Founder-Led Marketing (2025-03-31), Paid Media Benchmarks (2025-03-31, and its data is labeled 2024), and Inbound Buying (2025-03-17) — are all more than 12 months old. 8 of 11 customer success stories carry dates of 2025-10-01 or older (NFP: 2025-05-20, over 15 months). Only Bonterra, Clari, and dotCMS were updated in July 2026.
Why it matters: AI answer engines heavily favor recently updated content when citing sources for comparison and evaluation queries — 76.4% of ChatGPT's most-cited pages were updated within the last 30 days (ConvertMate, Q4 2025), and AI-cited content is 25.7% fresher on average than traditional organic results (Ahrefs, August 2025). Attribution and pipeline measurement are Refine Labs' core differentiators, and case studies are the proof content buyers ask AI about — stale dates push these pages out of the dominant citation window while competitor content gets cited instead.
Recommended fix: Establish a refresh cycle for the ~13 stale high-value pages: update the attribution/measurement flagship posts with current data and republish with new dateModified, and refresh the 8 older case studies (even a substantive results-update paragraph with a new date qualifies). Prioritize The Attribution Mirage and the Paid Media Benchmarks post, which still carries 2024 benchmark data.
What we found: Every page under /success-stories/ has the title tag "Refine Labs" (no client name, no outcome), no meta description, and zero Open Graph tags — unlike the rest of the site, where titles and descriptions are well-formed. The pages carry Review-type JSON-LD but the descriptive metadata layer is absent.
Why it matters: Case studies are the highest-value proof content for agency evaluation queries ("Refine Labs results", "Refine Labs reviews", "Directive vs Refine Labs"). A title of just "Refine Labs" gives crawlers and answer engines no signal about what each page demonstrates (e.g., "Clari: 67% lower acquisition cost"), suppressing both retrieval and citation quality for the exact queries where this content should win.
Recommended fix: Populate the case-study page template with unique title tags in the pattern "{Client} case study: {headline outcome} | Refine Labs", meta descriptions summarizing the before/after and key metrics, and standard OG tags. This is a single template fix in the CMS applied across all 11 pages.
What we found: Multiple H1s on several pages (megan-bowen: 5, paid-media-benchmarks blog: 6, content-creative-services: 2, chris-walker: 2, blog hub: 2); no H1 at all on roi-calculator, podcasts, videos, and the Firstup case study. On 12 of 20 inventoried blog posts the only H2 is the boilerplate "More from Refine Labs" — body section titles are rendered as styled text rather than heading elements, so posts like Paid Search Playbook and The Attribution Mirage expose no semantic structure despite having well-organized sections.
Why it matters: Answer engines use heading structure to segment pages into retrievable passages and to label extracted answers. A 2,300-word article whose only H2 is boilerplate is retrieved as one undifferentiated block, reducing the chance any specific claim gets extracted and cited. Recently updated posts (CFO Case for Brand, AI Traffic, LinkedIn Employee Profiles) show the correct pattern — descriptive H2s per section — confirming the template supports it.
Recommended fix: Enforce one H1 per page in the CMS templates; convert styled section titles in older blog posts to real H2/H3 elements (match the pattern used in the July 2026 posts); add H1s to the utility pages that lack them.
What we found: The /roi-calculator page renders a "Product comparison" section containing literal "Lorem ipsum dolor sit amet" and "Feature text goes here" placeholder copy. Blog post templates render "xx min read" where the reading time was never populated.
Why it matters: Placeholder text on a commercial page is extractable content — an answer engine summarizing the ROI calculator page can surface the lorem ipsum, and it signals low content hygiene to both crawlers and human evaluators arriving from AI referrals.
Recommended fix: Remove or complete the product comparison section on /roi-calculator; populate or remove the reading-time token in the blog template.
What we found: The Firstup success story is a stub (about 200 words; its "the problem / BEFORE / the work" sections render empty with only a pull quote). /marketing-maturity-assessment has ~115 words of extractable text around a quiz embed. The /podcasts and /videos hubs have under 300 words each and no H1.
Why it matters: Thin pages cannot be cited. The Firstup stub is the sharpest case: it sits alongside 10 substantive case studies, occupies the URL an answer engine would retrieve for "Refine Labs Firstup results", and gives it almost nothing to extract (content_depth 0.2 vs 0.7-0.9 for sibling pages).
Recommended fix: Complete the Firstup case study to match the sibling template (problem, work, results sections). Add descriptive supporting copy to the assessment, podcasts, and videos pages so each carries at least one self-contained extractable passage.
What we found: sitemap.xml lists 88 URLs with <loc> only — no lastmod, changefreq, or priority on any entry. The nav-linked /content-creative-services URL is absent from the sitemap (only its canonical target /creative-gallery is listed).
Why it matters: Without lastmod, crawlers cannot prioritize recently updated pages for recrawl, which delays how quickly the site's frequent content refreshes (many pages were updated in late July 2026) propagate into AI answer engines' indexes. Freshness investment the team is already making is partially invisible to crawlers.
Recommended fix: Configure the CMS sitemap generator to emit lastmod from each page's actual modified date. Verify the sitemap regenerates on publish.
What we found: robots.txt exists but contains only a Sitemap declaration — no User-agent rules of any kind. All seven AI-relevant crawlers (GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Googlebot, Bytespider) are not mentioned and therefore implicitly allowed.
Why it matters: Nothing is blocked — this is the correct outcome for AI visibility. But implicit allowance means any future robots.txt edit (e.g., a template migration or a security plugin default) could silently block AI crawlers with no one noticing, and the file cannot express intentional policy differences (e.g., allowing retrieval bots while opting out of training bots).
Recommended fix: Add explicit User-agent: blocks with Allow: / for the AI crawlers the company wants (at minimum GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot), making the allow decision durable and auditable.
The following items could not be assessed through our analysis method (rendered markdown). We recommend your engineering team verify these manually before the validation call.
What to check: JSON-LD is present and appropriately typed on most pages (BlogPosting on all blog posts, Organization/WebSite on the homepage, OfferCatalog on pricing, Service on The Vault, Review on case studies). This analysis parsed the raw HTML and confirmed the types, but did not validate required-field completeness (e.g., whether Review blocks carry itemReviewed/reviewRating, or BlogPosting blocks carry author/image as objects), and "Review" is an unusual type choice for case-study content where Article/BlogPosting is the conventional pattern.
Recommended action: Run the case-study, pricing, and Vault templates through Google's Rich Results Test / schema.org validator; fill missing required fields and consider Article-type markup (or Review with complete itemReviewed) on success stories.
Partial sample The analysis covered 49 of the 88 URLs in the sitemap (~56%), and 16 of the 49 analyzed pages — including all 10 product/commercial pages — carry no detectable date, so freshness could not be scored for them. The scores above are representative of the analyzed sample, not the full site; the undated product pages in particular should be verified manually.
Why now
• AI search adoption is accelerating — 87% of B2B software buyers say AI chatbots are changing how they research vendors, and half now start their research in a chatbot rather than Google (G2, October 2025).
• Early citations compound: domains that AI platforms learn to trust now get cited more frequently as training data accumulates.
• Competitors who establish GEO visibility first create a structural disadvantage for late movers — the answer slot they occupy is the one you have to displace.
• The B2B demand generation agency category is still early-innings in GEO optimization — acting now means competing against inaction, not against entrenched strategies.
The full audit will measure citation visibility across buyer queries in the B2B demand generation agency space — queries like "best agency to fix MQLs that never turn into pipeline," "our CAC has doubled — which B2B demand gen agency actually lowers it," and "Refine Labs vs Directive for B2B SaaS." You'll see exactly which queries return your competitors but not Refine Labs, and what it would take to appear in them — and because the Layer 1 fixes above will already be underway, your technical baseline improves before the audit even measures it.
45–60 minutes. We walk through this document together — you confirm, correct, and fill gaps. Every answer sharpens the query set before it runs.
We generate the full buyer query set from the validated knowledge graph and execute it across the selected AI platforms, capturing every response and citation.
Complete visibility analysis: where you're cited, where competitors win, and a three-layer action plan prioritized by what actually costs you citations.
Start now — no need to wait for the call Three engineering-side fixes can begin immediately: (1) fix the case-study page template so all 11 success stories get unique titles, meta descriptions, and OG tags — a 1–3 day single-template change; (2) configure the sitemap generator to emit lastmod timestamps for all 88 URLs — under a day; (3) remove the lorem ipsum placeholder content on /roi-calculator and add explicit Allow rules for AI crawlers (GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot) to robots.txt — under a day combined. These don't depend on the rest of the audit and will improve your baseline visibility before we even measure it.
Two jobs before we meet. The questions on the left require your judgment — no one knows your business better than you. The engineering tasks on the right don't require the call at all.