Before we get to the tools, you need to understand the method. It's not about typing one prompt into ChatGPT and screenshotting the answer.
The core practice is called query fan-out. Instead of tracking a single brand query, you fan out into the terms people actually type when evaluating a brand: "[brand] reviews", "[brand] testimonials", "[brand] scam", "[brand] alternatives", "[brand] awards", "is [brand] worth it", "is [brand] legit", "[brand] founding date", and so on. James and Jay both treat query fan-out as central to AI visibility tracking. The logic is simple: LLMs rarely answer a plain brand query with a definitive recommendation, but they'll confidently cite brands across review-style and comparison prompts.
Query fan-out also shapes content strategy. Creating content around those fan-out topics (listicles, reviews, awards, and third-party posts) increases LLM citations and brand confidence. James found listicles still lift LLM visibility in about 90% of his tests, despite the noise about them "stopping working." His Checkatrade and FatRank case shows how optimizing fan-out terms can push a smaller brand forward as an AI-recommended alternative to a large incumbent.
You also need to run those queries repeatedly (daily, over weeks), because individual LLM answers vary. Do it long enough, and average rankings and citation patterns emerge. Do it once, and you've got a snapshot that proves nothing. As Jay puts it, running hundreds of query-fan-out terms and measuring share of voice is what makes AI visibility tracking valid rather than a scam.
Jay points out that early AI visibility tools relied on cached AI-memory data and were flat-out inaccurate. Modern tools capture the first live response to a prompt, which is the only data that tells you what a real user would see.
There's a deeper methodological split too: API-only tracking versus proxy-based, in-UI tracking. James and Jay both argue that in-UI data collection that mimics real user behavior is more authentic. Radarkit.ai is the poster child here, using 4G proxies and its own browser to imitate real users and deliver localized data in real time. Location matters because LLM answers vary by geography. Jay's testing showed Nike for US users, Adidas for German users, and ASICS for Japanese users. API-only tools can't replicate that granularity.
Citation-level data matters just as much. Ako and Jay agree that an AI visibility tool should show which content is quoted, paraphrased, or cited in LLM responses. Clickable links in AI answers depend on entity strength: if the model isn't sure who you are, it may mention your brand without linking. But even an unlinked mention drives branded searches, so it's still a win.