Tag: Perplexity

  • How 215,000 Robot-Written Pages Tricked AI Into Recommending Software

    How 215,000 Robot-Written Pages Tricked AI Into Recommending Software

    When you ask an AI assistant for the best project management tool, it often pulls from a handful of websites. But a new investigation reveals that three of those sites are not what they seem—they’ve published over 215,000 pages of “best software” content, likely generated by bots, and AI systems like Perplexity treat them as trusted sources.

    This isn’t just a quirk of search algorithms. It’s a sign that the AI-powered web is vulnerable to a new kind of spam—one that doesn’t target Google rankings but targets the very systems that power AI answers. The result is a feedback loop where low-quality content gets elevated simply because it exists at massive scale.

    The Scale of the Problem

    Imagine a single editorial team trying to write genuinely useful “best software” articles. Each piece would require hands-on testing, expert opinions, and careful updates. A realistic operation might publish a few hundred per year. Yet three sites have collectively published 215,128 pages of this content—a number that would take a human team centuries to produce.

    This is the finding from a report by Trellner.com, titled “Manufactured Sources Behind AI Recommendations.” The report, which gained traction on Hacker News, exposes how these sites use programmatic SEO (pSEO) to generate pages at industrial scale. Each page is technically unique—different titles, different introductory paragraphs—but they all follow the same template, often scraping data from software vendor sites and wrapping it in boilerplate opinion.

    How Programmatic SEO Works

    Programmatic SEO is not new. It’s a technique where publishers use templates and databases to create thousands of pages targeting specific search queries. For example, a site might have a database of 500 software categories and 50 use cases, then generate a page for each combination: “best CRM for real estate agents,” “best CRM for nonprofits,” “best CRM for startups.” That’s 25,000 pages from just one niche.

    What makes this different from classic content spinning is that modern pSEO often uses real data. The pages might include accurate pricing tables, feature lists scraped from vendor sites, and even genuine user reviews. The problem is that the “opinion” content—the actual recommendation—is written by algorithms, not humans. A human editor never tests the software or forms a genuine opinion.

    The Trellner report identified three specific sites that have mastered this technique. While the report names them, the key takeaway is that they’re not obscure spam sites—they rank well and are cited by AI assistants. That’s because their pages perfectly match the kind of long-tail queries people type into AI tools.

    Why AI Systems Fall for It

    AI assistants like Perplexity use a technique called Retrieval-Augmented Generation (RAG). When you ask a question, the system retrieves relevant web pages, then uses them to craft an answer. The retrieval step is based on signals like keyword matching, domain authority, and link popularity—not on editorial quality.

    Content farms exploit this by producing pages that exactly match common AI prompts. If someone asks “best project management software for small teams,” the AI finds a page with that exact phrase in the title and content. The page looks authoritative because it has thousands of words, lists many options, and cites data from vendor sites.

    The result is a perverse incentive: instead of chasing Google rankings, spammers chase AI citations. The traffic from an AI citation isn’t a click on a search result—it’s being named in an AI answer, which drives users to visit the cited page. And with Perplexity’s revenue-sharing program, publishers get paid when their content is cited, creating a financial motive for manufacturing content specifically to be cited.

    The Feedback Loop

    Once an AI system cites a page, that citation can boost the page’s apparent authority. Other AI systems may see the page as a trusted source because it’s cited by Perplexity or ChatGPT. This creates a feedback loop where manufactured content gets elevated simply because AI systems reference each other’s sources.

    This is particularly damaging for legitimate publishers who invest in real editorial content. A journalist who spends weeks testing software and writing a nuanced review is competing against thousands of templated pages that can be generated overnight. The AI systems don’t distinguish between the two—they just see relevant keywords and high page counts.

    The problem is systemic. As one Hacker News commenter noted, if AI systems reward volume and keyword matching, publishers will optimize for that. It’s not just the fault of the three sites; it’s a flaw in how AI retrieval works.

    What AI Companies Say

    Perplexity and other AI companies face a quality control challenge. They cannot manually vet every source they cite. They rely on ranking signals—PageRank-like metrics, domain authority, freshness—that content farms can manipulate. A site that publishes 100,000 pages will naturally accrue a lot of internal links, which can boost its perceived authority.

    Perplexity’s likely defense is that they’re continuously improving their algorithms to detect low-quality content. But the Trellner report suggests the problem is structural. As long as AI systems rely on scalable signals, they will be vulnerable to scalable manipulation.

    The Broader Implications

    This story is not just about software recommendations. It’s about the integrity of AI-powered answers. If AI systems are supposed to provide trustworthy information, their citation infrastructure must be robust against gaming. Otherwise, they risk becoming a platform for spam—just like Google search results were in the early 2000s.

    The Trellner report is an example of independent investigation—a smaller outlet doing the kind of work that major tech journalism hasn’t yet covered systematically. It highlights a growing issue: the supply chain of AI information is being polluted by manufactured content.

    For users, the takeaway is to be skeptical of AI recommendations, especially for commercial queries. The software that an AI suggests may not be the best—it may just be the one with the most pages written about it.

    The discovery of 215,128 robot-written “best software” pages—and their prominence in AI citations—reveals a critical weakness in how AI systems gather information. As AI assistants become our primary gatekeepers to knowledge, the quality of their sources matters more than ever. Without better detection methods, the web risks being flooded with content designed not to inform, but to game the machines we trust to inform us.

    Summary

    • Three sites published 215,128 “best software” pages, likely generated via programmatic SEO, and AI assistants like Perplexity cite them.
    • Programmatic SEO uses templates and scraped data to create thousands of similar pages, targeting long-tail queries that match AI prompts.
    • AI systems retrieve sources based on keywords and authority signals, which content farms can manipulate at scale.
    • Perplexity’s revenue-sharing program creates a financial incentive to manufacture content for AI citations.
    • The problem is systemic, affecting the integrity of AI recommendations and crowding out legitimate publishers.

    FAQ

    Q: What is programmatic SEO?
    A: Programmatic SEO (pSEO) is a technique where publishers use templates and databases to automatically generate thousands of web pages. Each page is technically unique but follows a fixed structure, often targeting specific search queries to attract traffic.

    Q: How do AI assistants like Perplexity decide which sources to cite?
    A: They use retrieval-augmented generation (RAG), which pulls relevant web pages based on keyword matching and authority signals like domain age and link popularity. The process does not evaluate editorial quality.

    Q: Why are software recommendation pages particularly vulnerable to this type of spam?
    A: Software queries have high commercial intent (people often buy products after reading reviews), and the niche is data-rich—pricing and features can be scraped from vendor sites. This makes it easy to generate many pages with real data but fake opinions.

    Q: Does this affect only Perplexity, or other AI tools too?
    A: The report focuses on Perplexity, but any AI assistant that uses web retrieval—like ChatGPT with browsing or Google’s AI Overviews—can be vulnerable to similar tactics.

    Q: What can users do to avoid being misled by AI recommendations?
    A: Be skeptical of recommendations, especially for commercial products. Cross-check with multiple sources, look for human-authored reviews, and consider the possibility that the AI is citing content farms.

  • AI Search Market Share: What Business Intelligence Teams Need to Know

    AI Search Market Share: What Business Intelligence Teams Need to Know

    The search landscape is shifting under the feet of every business. Traditional link-list search the familiar ’10 blue links’ is being supplemented, and in some cases replaced, by AI-powered answer engines that synthesize information directly. For business intelligence (BI) teams, this isn’t just a tech trend; it’s a fundamental change in how customers find information, how advertising works, and how brand authority is measured.

    As of early 2025, Google still commands roughly 90% of the global search market, but AI-native search engines like Perplexity and ChatGPT Search are carving out a small but rapidly growing niche. More importantly, AI is being embedded into the incumbent’s product itself: Google AI Overviews now appear on a significant share of search results. This article breaks down the current market share data, the strategic positions of key players, and the practical implications for businesses that rely on search for growth and intelligence.

    Defining AI Search: More Than Just a New Engine

    AI search refers to search experiences that leverage large language models (LLMs), generative AI, and retrieval-augmented generation (RAG) to answer queries directly, synthesize information from multiple sources, and provide conversational results. This is distinct from traditional search engines like Google, Bing, or DuckDuckGo, which primarily deliver a list of links ranked by relevance.

    Key players in the AI search space include:

    • ChatGPT Search (OpenAI), launched in October 2024, which integrates real-time web access into the ChatGPT interface.
    • Perplexity, a purpose-built ‘answer engine’ that provides cited, conversational responses.
    • Google AI Overviews and Gemini, which embed AI-generated answers directly into Google’s search results.
    • Microsoft Bing Copilot, which uses GPT-4 to power conversational search on Bing.
    • You.com, Brave Search AI, and emerging entrants like Meta AI search and Amazon Rufus for shopping.

    Market Share: The Numbers Behind the Hype

    Despite the buzz, AI-native search engines still hold a tiny slice of the overall search pie. Perplexity, ChatGPT Search, and You.com collectively account for less than 1–2% of global search queries. Google remains the dominant force at ~90%, with Bing at 3–4%, and a long tail of others making up the remainder.

    Yet that small percentage masks significant momentum in specific segments. Perplexity, the leading standalone AI search engine, has grown to an estimated 15–20 million monthly active users, with reported annualized revenue of ~$50 million in late 2024. ChatGPT Search, leveraging OpenAI’s existing base of over 200 million weekly active users, has seen rapid adoption, though OpenAI does not disclose search-specific usage figures.

    Perhaps the most impactful shift is within Google itself. AI Overviews are now estimated to appear on 30–80% of Google search results pages, depending on geography and query type. This means AI is not just a separate market—it’s becoming the default experience for many users on the world’s most-used search engine.

    Why Business Intelligence Teams Should Care

    The rise of AI search has profound implications for how businesses track, attribute, and optimize their online presence.

    1. Ad Spend and the Zero-Click Problem

    If AI search answers a query directly, users may never click through to a website. This ‘zero-click’ behavior breaks traditional pay-per-click (PPC) economics, where advertisers pay for each visit. Marketers need to understand where these zero-click answers dominate and adjust their strategies accordingly. For example, if a user asks ‘best CRM for small business’ and gets a synthesized answer from ChatGPT Search, the opportunity for a traditional ad impression may vanish.

    2. Data and Attribution Shifts

    AI search engines cite sources differently than traditional link lists, and referral traffic patterns are changing. BI teams must develop new tracking methods—such as AI-specific UTM parameters and server-side tracking—to accurately measure the impact of AI search on their web traffic. Monitoring which AI engines surface a brand (or a competitor) first is becoming a key performance indicator.

    3. Competitive Intelligence

    Tracking your brand’s visibility in AI search results—and your competitors’—is a new frontier for competitive intelligence. Are you cited as a source in Perplexity’s answers? Does ChatGPT Search mention your product in a comparison? These metrics can serve as leading indicators of brand authority in the AI era.

    4. Enterprise Search as a Related Market

    Internal AI search tools like Glean and Microsoft Copilot for M365 are a separate but related market, growing at over 30% CAGR. Businesses are deploying these tools to help employees find information faster, and BI teams may need to integrate data from these platforms into their analytics.

    The Incumbent’s View: Google’s Defense

    Google argues that AI Overviews are an enhancement, not a replacement, and points to high user satisfaction metrics. The company’s counter-strategy includes deep integration of its Gemini model, an AI Mode for search, and maintaining default search deals (like the one with Apple) that keep its market share locked in.

    However, Google faces a real risk: AI Overviews reduce click-through rates to publishers, potentially undermining the open web ecosystem that feeds its index. If publishers see less traffic, they may produce less content, which could degrade the quality of Google’s search results over time. Antitrust remedies from the U.S. DOJ case could also force changes to Google’s default search deals, opening distribution windows for competitors.

    The Challenger’s View: Perplexity and OpenAI

    Perplexity positions itself as an ‘answer engine’ with citation transparency, appealing to researchers and business professionals who need verifiable sources. It was the first to launch a standalone AI search product and has built a loyal following among tech-savvy users.

    OpenAI, on the other hand, sees search as a feature within a broader assistant ecosystem, not a standalone product. Its distribution advantage is enormous: over 200 million people already use ChatGPT. The company is also exploring advertising, which would directly compete with Google’s core ad business.

    Both challengers face high compute costs, and their monetization models—subscriptions and nascent ads—are unproven at scale. Perplexity launched ads in 2024, and OpenAI is testing advertising in ChatGPT, but it’s unclear how much revenue these can generate.

    The Publisher’s Dilemma: To Block or to License

    AI search reduces referral traffic to publishers, prompting many to block AI crawlers (e.g., The New York Times, Reuters) or negotiate licensing deals. However, there’s a counter-narrative: AI search can drive high-intent traffic if a brand is cited as a source. In this new paradigm, ‘being the answer’ is the new SEO. For BI teams, tracking AI-citation share can be a leading indicator of brand authority and future organic traffic.

    What This Means for Your BI Strategy

    1. Monitor AI search visibility: Set up tracking to see how often your brand appears in AI search results and from which engines.
    2. Adjust attribution models: Incorporate AI search referrals into your analytics, using custom parameters and server-side tracking to capture data accurately.
    3. Re-evaluate SEO: Traditional keyword optimization is still relevant, but focus on creating content that AI engines are likely to cite as authoritative sources.
    4. Track ad performance: Understand where zero-click answers dominate and adjust your PPC campaigns accordingly.
    5. Stay agile: The market is evolving rapidly—what’s true today may change next quarter. Keep an eye on new entrants and shifts in user behavior.

    AI search is not a distant future—it’s happening now. While AI-native engines still hold a small market share, the integration of AI into Google’s core product means that AI-generated answers are already a significant part of the search experience. For business intelligence teams, the imperative is clear: adapt your tracking, re-evaluate your strategies, and start treating AI search visibility as a critical metric. The search landscape is changing, and those who understand the shift will be better positioned to thrive in it.

    Summary

    • AI-native search engines (Perplexity, ChatGPT Search, You.com) hold <1–2% of global search queries, but Google AI Overviews appear on 30–80% of search results pages.
    • Perplexity leads standalone AI search with 15–20 million monthly active users and ~$50M annualized revenue.
    • ChatGPT Search leverages OpenAI’s 200M+ weekly users, but search-specific usage is undisclosed.
    • AI search disrupts traditional PPC models due to zero-click answers, requiring new tracking and attribution methods.
    • Businesses should monitor AI-citation share as a leading indicator of brand authority.

    FAQ

    Q: What is AI search?
    A: AI search uses large language models and generative AI to answer queries directly with synthesized, conversational results, rather than just providing a list of links.

    Q: How big is the AI search market?
    A: AI-native search engines currently hold less than 1–2% of global search queries, but Google AI Overviews—which are AI-generated—appear on a significant share of Google’s results pages.

    Q: Who are the main players in AI search?
    A: Key players include Perplexity, ChatGPT Search (OpenAI), Google AI Overviews/Gemini, Microsoft Bing Copilot, and You.com.

    Q: How does AI search affect advertising?
    A: AI search can lead to zero-click answers, where users don’t click through to websites, which disrupts traditional pay-per-click advertising models.

    Q: How can businesses track their performance in AI search?
    A: Businesses can use AI-specific UTM parameters, server-side tracking, and monitor AI-citation share to measure visibility and referral traffic from AI search engines.

  • Generative Engine Optimization: Making Your Content Visible to AI Answers

    Generative Engine Optimization: Making Your Content Visible to AI Answers

    When you ask ChatGPT or Perplexity a question, the answer you get is often a synthesized paragraph built from a handful of sources. The AI doesn’t show you a list of blue links; it gives you a summary with citations. For website owners, this changes everything. If your content isn’t one of those cited sources, you’re invisible to a growing audience that gets answers without ever clicking through.

    This shift from search to answers is what Generative Engine Optimization (GEO) addresses. GEO is the practice of making your content more likely to be cited, summarized, or recommended by AI-powered engines. It’s not a replacement for traditional SEO it’s a new layer focused on how AI models retrieve and trust information.

    What Exactly Is GEO?

    Generative Engine Optimization (GEO) is the art and science of optimizing content so that AI answer engines like ChatGPT, Perplexity, Google AI Overviews, and Bing Copilot pick it up and feature it in their responses. Traditional SEO targets search engine crawlers and ranking algorithms to win a top spot on a results page. GEO targets large language models (LLMs) and their retrieval systems to win a citation in a synthesized answer.

    The term was coined in a February 2024 paper by researchers at Princeton, Georgia Tech, and IIT Delhi. The paper, “Generative Engine Optimization: A New Paradigm for Content Optimization,” showed that adding quantitative, statistical, and factual language to content increased its visibility in AI-generated answers by up to 40%.

    How AI Engines Choose Sources

    Most AI answer engines use a technique called Retrieval-Augmented Generation (RAG). Here’s how it works in plain terms: when you ask a question, the engine first searches through a vast index of documents, looking for ones that are semantically similar to your query. That means it’s not just matching keywords it’s understanding meaning. Then, it feeds the top few documents to a large language model, which reads them and generates a coherent answer.

    So which documents get picked? Research and observation suggest that AI engines favor sources that:

    • Are semantically close to the query’s meaning, not just its keywords.
    • Contain explicit, extractable facts clear numbers, dates, names, and statements.
    • Come from domains that appear authoritative (high domain trust, recognized expertise).
    • Are structured in a way that makes information easy to pull out lists, tables, clear headings, and concise paragraphs.

    There’s also a training data bias: content that appears frequently in a model’s training data (often older, high-authority content) has an inherent advantage. If your content has been around for years and is widely referenced, the AI is more likely to recall it.

    GEO vs. SEO: What’s Different?

    To understand GEO, it helps to contrast it with SEO. SEO focuses on keywords, backlinks, and technical site structure. Its goal is to get a page to rank in the top 10 results on a search engine results page (SERP). GEO focuses on entity clarity, structured data, quotable statistics, and semantic authority. Its goal is to be one of the 3–5 sources an AI cites in a synthesized answer.

    Here’s a concrete example. In SEO, you might write an article about “best running shoes” and optimize it for the keyword phrase. In GEO, you’d also include a clear list of top shoes with specific features, a table comparing prices, and a concise summary at the top. When an AI engine retrieves content to answer “what are the best running shoes?”, it can easily pull facts from your structured list and citation.

    Why GEO Matters Now

    The shift from search to answers is happening faster than many expected. In May 2024, Google launched AI Overviews, which appear at the top of billions of queries. Bing integrated GPT-4 back in 2023. Dedicated answer engines like Perplexity are growing rapidly, especially among younger users who prefer direct answers over link lists.

    This matters because of the “zero-click” dynamic. In traditional search, users click through to websites. With AI answers, users may never leave the results page. For publishers, being cited is now a primary traffic driver. If your content isn’t cited, you’re missing out on a growing share of user attention.

    Practical GEO Tactics You Can Use

    So how do you make your content AI-friendly? Based on early research and practitioner experience, here are concrete steps:

    1. Add clear, quotable statistics. The GEO paper found that adding quantitative language boosts visibility. If you’re making a claim, back it with a number. Instead of “many companies use AI,” say “62% of companies report using AI in some form.”
    2. Structure your content with headings and lists. AI engines love extractable information. Use H2 and H3 headings to break up your content, and use bullet points or numbered lists for key facts. This makes it easy for an LLM to pick out the exact sentence it needs.
    3. Write a concise summary at the top. A TL;DR section or a “Key Takeaways” box helps AI engines quickly grasp what your page is about. It also improves user experience.
    4. Use structured data (schema markup). While GEO is still evolving, structured data helps AI understand your content’s entities. Implement schema types like Article, FAQ, or Product to give clear signals about what your page covers.
    5. Focus on entity clarity. Make sure your content clearly identifies the main entities—people, places, products, concepts—and their relationships. Use consistent names and avoid ambiguous references.
    6. Cite authoritative sources. When you reference external data, link to high-authority sources. This builds trust and makes your content more likely to be considered authoritative itself.
    7. Include quotes and expert opinions. Research shows that AI engines often cite content with direct quotes. If you have an expert quote, include it verbatim.
    8. Keep content fresh. AI models update their training data and retrieval indexes. Regularly updating your content keeps it relevant and increases the chances it will be cited.

    The Skeptical View: Is GEO a Moving Target?

    Not everyone is convinced GEO is a stable discipline. Skeptics point out that LLM behavior changes with each model update. A tactic that works today might not work next year. This is a valid concern. Search engines also change their algorithms constantly, yet SEO has evolved into a mature practice. GEO is likely to follow a similar path, but it’s still early.

    Another concern is “citation without traffic.” Being cited in an AI answer doesn’t necessarily mean users click through to your site. The answer itself might satisfy the query completely. Some argue that GEO should focus on brand visibility rather than direct traffic. If your brand is cited as an authority, that builds trust even if people don’t click immediately.

    Where GEO Is Heading

    The field is young, but it’s growing fast. Major SEO agencies now have GEO practice areas. Academic research is continuing, with follow-up studies on citation behavior. And AI platforms are experimenting with ad placements inside AI answers, creating a new advertising surface that could compete with organic citations.

    For website owners, the message is clear: start optimizing for AI now. The strategies are not radically different from good content practices—clarity, authority, and structure—but they’re tailored to the way AI consumes information. By making your content more citable, you position yourself to remain visible in the new answer economy.

    Generative Engine Optimization is not a fad; it’s a response to a fundamental shift in how people get information. As AI answers become the default, the ability to be cited by these engines will determine your online visibility. The good news is that GEO builds on solid content practices: be clear, be specific, be structured. By adopting GEO tactics now, you’re not just optimizing for algorithms—you’re ensuring that when someone asks an AI a question, your expertise is part of the answer.

    Summary

    • GEO (Generative Engine Optimization) optimizes content for AI answer engines like ChatGPT, Perplexity, and Google AI Overviews.
    • Unlike SEO, which targets keyword ranking, GEO focuses on being cited in AI-synthesized answers.
    • Key tactics include adding statistics, using structured data, writing clear summaries, and maintaining entity clarity.
    • The term was coined in a 2024 academic paper that showed a 40% boost in visibility from quantitative language.
    • GEO is still evolving, but early adoption can help you stay visible as AI answers grow.

    FAQ

    Q: What is the difference between SEO and GEO?
    A: SEO optimizes for search engine crawlers to rank high in link results; GEO optimizes for AI models to be cited in generated answers. GEO focuses on semantic clarity, structured data, and quotable facts.

    Q: How do AI engines decide which sources to cite?
    A: They use retrieval-augmented generation (RAG), which first finds documents similar to the query, then feeds them to an LLM. They favor sources that are semantically relevant, contain explicit facts, come from authoritative domains, and are well-structured.

    Q: Does GEO require completely new content?
    A: Not necessarily. You can adapt existing content by adding summaries, statistics, and better structure. The goal is to make information easy for AI to extract.

    Q: Is GEO worth it if AI answers don’t send clicks?
    A: Yes, for brand visibility. Being cited positions you as an authority, even if users don’t click through immediately. Over time, this can lead to direct visits and trust.

    Q: What’s the biggest challenge in GEO?
    A: The field changes quickly as AI models update. Tactics that work today may need adjustment tomorrow. Staying informed and adapting is key.