Blog

We Ran 1,296 AI Searches Across DFW. Here Is the Engine We Built to Do It.

The full methodology behind Vimina's DFW study: a prompt model weighted against a database of millions of real queries, repeated runs to measure answer churn, controlled collection, and an auditable scoring model. Written by the engineers who built it.

Most agencies selling AI visibility will tell you what engines want. Almost none of them have measured it. We are engineers, so before Vimina sold anything, we built an instrument. One of us is finishing a master's in AI at Boston University. The other works at Columbia University. Our advisor holds a PhD in data analytics. What follows is the actual methodology, with the numbers, the constants, and the parts that went wrong. If you want the one-paragraph version: we ran 1,944 collection jobs across two AI answer surfaces and eight DFW cities, audited 10,279 cited sources by hand where the pipeline could not, and scored four sectors with a model whose every weight is written down.

The question, stated precisely

When a consumer in a specific DFW city asks an AI engine a high-intent local service question, which sources does the answer cite, how stable is that answer over time, and what fraction of it can a business legitimately influence? Vague versions of this question produce vague agencies. The precise version produces a measurement problem, and measurement problems have engineering answers.

Finding the right prompts, not guessing them

Garbage prompts, garbage study. We did not sit in a room inventing what we imagined people type. The prompt model has two layers. The first is an intent taxonomy: seven consumer intent types, from hyper-local (collision repair shop near me in Arlington) through how-to, comparative, multi-condition, opinion, and experimental, each carrying an explicit weight in the model. Hyper-local carries 0.30 because city-named and near-me queries dominate real local service demand. Every weight lives in a config file, not in someone's head.

The second layer grounds those templates in observed demand. We weight and expand queries against a commercial search database that tracks real query volume, related queries, and intent classification across millions of keywords. Our production scanner formalizes this as a leverage score: volume, commercial value, intent fit, and opportunity multiplied together, each factor between zero and one. Multiplication is deliberate. A query with huge volume but zero commercial intent scores near zero, and a query nobody types scores near zero, no matter how winnable. The sampling budget flows to prompts where all four factors are simultaneously nonzero. This is the difference between sampling what matters and sampling what was easy to think of.

For the DFW study this produced 140 seed templates across four sectors, geo-expanded by intent: hyper-local templates ran across all eight cities, multi-condition across four, everything else across the two anchor cities of Dallas and Fort Worth. That yields 280 concrete query units, and with engines and repeats applied, 1,944 collection jobs.

Why we ran the same query more than once

A single answer tells you who won today. It cannot tell you whether the ranking is settled or still in motion, and for anyone trying to enter the answer set, that second property is the entire game. So the high-stakes intent types, hyper-local and comparative, were run three times each. We then measure churn: for each pair of answers to the same prompt, take the two sets of cited domains and compute how much they disagree, using Jaccard distance. Zero means identical citations, one means no overlap at all. We compute this two ways, across repeated identical runs for temporal churn, and across city variants of the same template for geographic churn, blended 60/40. One edge case matters enough to name: two answers that both cite nothing are consistently sparse, not volatile, so they count as zero churn. Confusing those two states would overstate movability exactly where the engine is most starved.

High churn means the engine has not converged on a canonical source set, which means the ranking is movable, which means a well-structured entrant can take a slot. This number is why we can say a sector is open rather than feel that it is.

What we held constant

  • One model, one configuration: ChatGPT answers came through the OpenAI Responses API with the web search tool enabled, same model for every job, so differences between answers are differences in the answers, not in our client settings.
  • Pinned locations: every Google AI Mode call carries an explicit city-level locale, down to fixing the one city that resolves ambiguously (McKinney is pinned to Collin County). No answer was allowed to drift to a generic US result.
  • Immutable templates: prompts are generated by code from the seed matrix. Nobody rephrased anything mid-run.
  • Content-addressed caching: every raw response is cached under a key built from engine, query, and run number. The pipeline is resumable, never double-spends, and every number in the report can be traced back to a stored raw response.
  • A hard budget cap with a dry-run cost meter: the run aborts if projected spend exceeds the cap. Discipline is cheaper than apologies.

The scoring model, in plain English

Each sector gets an opportunity score built from three multiplied factors: how movable the answers are, how much competitive headroom exists, and what one customer is worth. Multiplication again, for the same reason as before: if any factor is near zero, the opportunity is near zero, and adding would hide that. The movability factor is itself a weighted blend of four measured signals. Openness: one minus the concentration of citations, so a Wikipedia-style monopoly scores near zero and a flat field of many domains scores near one. Volatility: the churn measure above. Sparsity: how few distinct sources answers cite, plus the fraction of answers that cite nothing at all. And incumbent weakness: we crawl the top cited competitors and check, mechanically, whether they publish an llms.txt file, whether they ship structured data, and whether their content is fresh. Weak incumbents are overtakeable incumbents, and you can measure weak.

Lead value is not a guess either. It is average ticket times urgency times an insurance factor times close rate, per sector. Restoration prices at six thousand dollars a job with near-total urgency, which is why it can top the ranking despite junk removal generating more raw citations.

The audit nobody wanted to do

After collection, 89% of cited domains were unknown to our classifier, which would have made the capturability numbers meaningless. So we reviewed the long tail: 3,045 unknown domains covering 8,900 mentions. Nearly all turned out to be businesses' own websites, which are exactly the sources a business controls. Every promotion is written to an audit file with the reasoning attached. The handful of exceptions got classified explicitly: two aggregators reclassified as claimable directories, one shared site-builder domain scored as semi because no single business controls it. Final mix: 97.2% of citations capturable, 1.9% semi, 0.9% locked.

What the instrument found

  • Two engines, two playbooks. Claimable directory profiles were 15.6% of Google AI Mode citations and 0.3% of ChatGPT citations. ChatGPT cites businesses' own websites almost exclusively. The work required to win each surface barely overlaps.
  • ChatGPT answered with no cited source at all 8.5% of the time. Google AI Mode: never. Uncited answers are unclaimed answers.
  • Roughly 98% of all cited sources are ones a business can own or claim. The game is winnable, and now that is a measurement, not a slogan.
  • Restoration ranked first at 0.718 on our scale, on lead value and fragmentation. Home services had the most scattered answers we measured, which makes it the least settled field.

What went wrong, in print

Two things, and we publish both. First, our Google AI Overview collector had a bug: it never made the required second async fetch, so 82% of those scans came back empty and the raw content is unrecoverable without a re-run. We dropped that surface and used Google AI Mode, which was captured completely, as the Google signal. The substitution moved the final scorecard by less than 0.01. Second, hyper-local prompts ended up as 74% of run-cells, so comparative and how-to intents are under-sampled and the study is best read as a hyper-local study. A vendor who tells you their data has no asterisks is telling you they have not looked.

Where the engine went next

The DFW study was generation one. The production scanner that came out of it treats every AI engine as something you observe through surfaces, samples each observation with a confidence interval, and only reports a claim when independent surfaces agree. It refuses, by design, to pool a proxy measurement with a higher-fidelity one into a single number, and its reports say AI-search visibility rather than pretending a proxy is the consumer app. That restraint is also an engineering decision. The literature this all stands on is public: the generative engine optimization work presented at KDD 2024 measured average visibility gains around 40% from content changes, and up to 115% for lower-ranked sites, which is the published basis for believing this work compounds for challengers.

This is the same instrument we run monthly for clients. Measure, change what the engines read, measure again. If a vendor cannot show you the second measurement, they are selling you the first one forever.