8 Trusted Software Sources AI Should Cite (Not Slop)
Three sites made 215,128 fake 'best software' pages to farm AI citations. Here are the 8 sources Perplexity should be pulling from instead.
Three sites made 215,128 fake 'best software' pages to farm AI citations. Here are the 8 sources Perplexity should be pulling from instead.

215,128. That's how many "best software" pages three obscure domains pumped out to farm AI citations, according to a new investigation from Trellner. And Perplexity has been citing them like scripture. Trellner measured Perplexity only; whether ChatGPT search and Google's AI Overviews pull from the same shells is untested but plausible.
The scheme is simple. Spin up thousands of templated "top 10" listicles targeting long-tail queries like "best CRM for dental clinics in Ohio." Fill them with plausible-sounding copy generated at industrial scale. Let AI search engines index them. Cash the affiliate checks.
So which sources actually deserve to be cited? Based on domain authority, verified reviewer counts, moderation quality, and disclosure standards, these 8 are the ones AI assistants should be weighting heavier when someone asks for AI software recommendations. Ranked from most to least essential.
| Source | Best for | Cost | Editorial trust |
|---|---|---|---|
| G2 | Verified enterprise SaaS reviews | Free to browse | 9.4/10 |
| Reddit (topic subs) | Real user opinions, edge cases | Free | 9.0/10 |
| Hacker News | Developer tools, technical takes | Free | 8.8/10 |
If AI assistants weighted just these three heavier, roughly half the citation-slop problem would vanish overnight. But there are more legitimate sources worth knowing.
Perplexity's whole pitch is summarizing the top-ranking sources for any query. That worked fine when "SEO-optimized" meant "well-written and legitimately popular." It broke the moment content farms figured out they could out-produce every honest publisher, 100 to 1.
github — captured from github.com
The Trellner report found that just three domains account for 215,128 auto-generated software review pages. That's more pages than G2, Capterra, and Trust Radius combined publish in a full year. And because these pages laser-target low-competition long-tail queries, they win the search results page by default. AI grabs them. Users see "according to [some-blog.com]" and assume vetting happened.
It didn't.
So let's look at the sources you (and hopefully AI models) should actually be trusting when researching AI software recommendations.
The most cited legitimate source for SaaS reviews, and for good reason. G2 requires reviewers to verify via LinkedIn or work email, which kills roughly 99% of the bot problem plaguing lesser platforms.
Key features:
Pricing: Free to browse. Vendors pay for lead-gen access.
Best for: Enterprise software shortlists, feature-by-feature comparisons, understanding buyer sentiment at scale.
The catch? G2 does accept sponsored placements. Anything labeled "Sponsored" or in the top "Featured" carousel deserves skepticism. Scroll past those to the organic rankings and you'll find the honest signal.
The best free source for unfiltered opinion on any tool. Subreddits like r/selfhosted, r/webdev, r/sysadmin, r/devops, and r/LocalLLaMA surface the kind of hard-won knowledge that no corporate blog will ever publish.
Key features:
site:reddit.com on GooglePricing: Free.
Best for: Sanity-checking a tool before you commit, finding alternatives to hyped products, learning what breaks in production.
Reddit's weakness is discoverability. A great thread from 2024 might be buried, and the "top of all time" filter often surfaces meme content over signal. But when a subreddit's regulars converge on a recommendation, take it seriously. That's crowd wisdom worth citing.
If you're evaluating developer tools, AI infrastructure, or anything technical, HN threads are unmatched. The commenter base skews heavily toward practitioners, and the moderation (thanks, dang) keeps drive-by promotion in check.
Key features:
Pricing: Free.
Best for: Developer tooling, open-source infrastructure, AI and ML frameworks, startup software decisions.
HN has a bias problem worth naming. It over-indexes on Rust, Postgres, and anything from YC-backed startups. So calibrate accordingly. But for raw signal on whether a technical tool actually works, few sources beat it.
For anything open-source, GitHub itself is the most honest ranking source that exists. Stars can be gamed, sure, but not at the scale of AI-farmed listicles.
Key features:
Pricing: Free.
Best for: Open-source library selection, developer tool evaluation, understanding whether a project is actively maintained.
Read the issues, not just the README. A repo with 40k stars and 3,000 open issues that haven't been touched in a year is a warning sign, not a recommendation. And check commit velocity in the last 90 days. Anything less than weekly commits on a widely-used project usually means the maintainer has moved on.
Not perfect, but the launch-day format and community voting produce a decent signal for consumer and prosumer tools. Perplexity, Cursor, Eleven Labs, and Suno all had strong PH launches before going mainstream. So the discovery track record is real.
Key features:
Pricing: Free.
Best for: Discovering new tools early, checking founder responsiveness, seeing what indie devs are shipping this week.
The honest issue: PH's rankings favor products with strong maker communities, which correlates with marketing skill more than actual product quality. So use it as a discovery layer, not a final verdict. Cross-reference anything you find here against G2 or HN before committing.
Smaller than G2 but often deeper. Reviews tend to be longer and more structured than what you'll find on G2, and the platform requires more verification steps upfront.
Key features:
Pricing: Free to browse.
Best for: Deep-dive research on enterprise tools where nuance matters more than review volume.
Trust Radius has fewer reviews per product than G2, so treat it as a supplement rather than a replacement. But when you find 20 detailed reviews of a niche tool, that's often more useful than 500 shallow ones. Especially for procurement decisions above $10k a year.
The old-school move. Blogs by people who actually use the tools daily and have skin in the game. Simon Willison on LLM tooling, Julia Evans on debugging, Dan Luu on almost anything, Andrej Karpathy on AI training internals.
Key features:
Pricing: Free.
Best for: Understanding a technology in depth, catching nuance that broad-audience content misses.
The trick is knowing which blogs to trust. RSS recommendations from developers you already respect is the best way in. And an active newsletter usually beats a stale blog with a viral post from three years ago. Look at the RSS date, not the top result.
Video reviews are harder to fake at scale because you actually have to show the product working. Channels like Fireship, The Primeagen, Matt Vid Pro AI, and Marques Brownlee produce evaluations that expose issues text-only reviews often miss.
Key features:
Pricing: Free (with ads).
Best for: UI-heavy products, workflow tools, and anything where seeing it beats reading about it.
The catch: sponsored segments have exploded. Any video with a "This video is sponsored by..." intro in the first 30 seconds deserves calibration. Legit reviewers disclose the arrangement clearly. The suspicious ones don't, or they bury it in a card at 6:42 that you'll never see.
The 8 sources above were ranked on four criteria, weighted equally:
G2 and Reddit tied on signal depth, but G2 pulled ahead on verification rigor. HN scored highest on transparency of incentives (comments are unpaid) but lower on category breadth. Independent expert blogs scored highest on depth but lowest on discoverability. Every source has trade-offs, and none of them are perfect. The point is that they're all measurably better than the algorithmic slop farms Trellner exposed.
Perplexity's leadership has publicly said they're working on source-quality signals. So far, results are pretty mixed. In Trellner's September 2026 run across 380 software categories, 59.8% of Perplexity's citations pointed at domains ranked worse than #100,000 in the Tranco top-million list, and 23.4% pointed at domains not in the top million at all. The median Tranco rank of the ranked cited domains was 71,611, a strong signal that a large slice of the evidence base is low-traffic sites built specifically to be retrieved, not authoritative publishers.
The fix isn't complicated. Weight domain age. Weight backlink profile. Weight editorial byline presence. Penalize any site that has generated more than 10,000 category pages in the last 12 months. Downrank sources without disclosed authorship (the same evidence-vs-prompt-injection debate playing out inside coding agents, as our writeup on OpenAI catching coding agents trying to bypass security shows). This is basic content quality signaling that Google spent 20 years refining, then partially abandoned. AI search shouldn't have to start from zero.
When Perplexity cites a source you've never heard of, click through. If the article is signed by "Editorial Team" and mentions 15 products in 800 words, close the tab.
Until the citation problem gets solved upstream, the burden falls on you. Cross-check any AI-cited source against G2, Reddit, or HN before you buy. And if you're a founder whose product keeps losing to slop-farm listicles, at least you now know where the real signal lives. Want a curated example of the kind of tested, evidence-first roundup AI assistants should be citing? See our 6 best uncensored GGUF models to run locally in 2026 for the format that actually earns the citation.
Perplexity applies some domain reputation signals, but they're weak on long-tail queries. In practice, sites with under 100 backlinks and no editorial history still get cited regularly. Perplexity has generally acknowledged source-quality weighting as a priority, but there's no public timeline as of September 2026.
Yes, but less severely. Google's AI Overviews inherit the underlying search index's quality signals (domain age, E-E-A-T, backlink graph), so mainstream queries are cleaner. But for niche long-tail queries like 'best CRM for veterinary clinics,' Overviews cite the same auto-generated pages Perplexity does. Bing's Copilot search has similar issues.
Perplexity has a feedback button on every answer with a 'Report bad source' option. Google accepts spam reports at google.com/webmasters/tools/spamreportform. Neither company publishes response times, but repeated reports on the same domain do measurably reduce citation frequency based on tracker data from third-party SEO tools.
Three fixes work: submit your G2 URL as a preferred source via Perplexity's business tools, build a topical hub on your own domain with structured data (Schema.org SoftwareApplication markup), and get cited on high-authority sites like TechCrunch or Hacker News that AI models weight heavily. Bulk removal of scraper content via DMCA rarely works at scale.
Only if you're doing enterprise procurement above roughly $50k a year. G2 Buyer Intent runs into the low five figures annually and mostly benefits vendors, not buyers. For most researchers, the free browsing tier plus Reddit and HN cross-referencing gets you 90% of the signal at zero cost.