The headline: bot and AI traffic overtaking human traffic
You've probably seen the claim: bots now account for more internet traffic than humans do. It shows up in tech press roundups, in "state of the internet" reports, and increasingly in the People Also Ask box when anyone searches for how AI is changing the web.
The claim itself isn't new. Bot traffic has outpaced human traffic in various reports for years, long before generative AI existed, driven by search engine crawlers, uptime monitors, scrapers, and spam. What's changed recently is the composition: a growing share of that bot traffic now comes from AI crawlers indexing content for chatbots and AI answer engines, on top of the traditional search crawlers, scrapers, and monitoring bots that were already there. We're not going to attach a specific percentage to this without a source we can verify, and you shouldn't trust any blog post that hands you a precise number without citing where it came from. The direction of the trend is well documented. The exact ratio depends on who measured it, when, and how they defined "bot."
For a self-hosted blog or a small business site, this isn't an abstract industry statistic. It's a data quality problem sitting inside your analytics dashboard right now. If a meaningful share of your "traffic" is crawlers, scripts, and AI agents fetching pages, then pageview counts, session counts, and even some engagement metrics are measuring something other than what you think they're measuring.
In short: a rising share of requests hitting any public website are automated, not human, and that share includes a growing set of AI crawlers on top of the search and monitoring bots that were already there. If you read your analytics as a direct measure of human interest without separating bot traffic first, you are measuring the wrong thing.
Why this matters more for a small blog than a big media property
A large media site has enough raw traffic volume that bot noise is a rounding error lost in the aggregate, and they usually have dedicated infrastructure (bot management services, server log analysis, engineering time) to filter it out. A small blog doesn't have that cushion, and usually doesn't have that infrastructure either.
If you publish a post and your dashboard shows 40 visits in the first day, and even a handful of those are crawlers, scrapers, or AI agents fetching the page for indexing, the distortion is proportionally huge. A big site with 400,000 visits absorbs 40 bot hits without anyone noticing. A blog with 40 total visits can have its entire signal replaced by noise.
This matters for three concrete reasons:
You make decisions off small numbers. When you're deciding whether a topic resonated, whether a title worked, or whether to write a follow-up post, you're often looking at differences of single digits or low double digits between posts. Bot noise at that scale can flip your conclusion.
You don't have a baseline to compare against. A large publisher can compare this month's bot ratio to last month's and spot the delta. A new blog has no history, so every session looks equally legitimate by default.
Crawl traffic to a blog is going up, not down, as AI answer engines expand. Beyond the crawlers search engines have always run, a widening set of AI systems fetch pages to build indexes, ground answers, or check for updates. That's a second, mostly new source of automated requests layered onto a small site's traffic, on top of the bots that were already there.
None of this means your blog doesn't have real readers. It means the ratio of signal to noise in your raw traffic numbers is worse than it looks, and it's worse specifically because your total volume is small.
What your analytics should actually separate out
Most default analytics setups report one number: visits, or sessions, or pageviews, as if it's a single clean signal. It isn't. To get a number you can actually make decisions from, split it into at least these categories:
Known bots and crawlers. Search engine crawlers (Googlebot, Bingbot) and known AI crawlers identify themselves by user agent string. A decent analytics setup filters these out of your human-traffic count by default, or at minimum tags them separately so you can see them without them polluting your headline number.
Unidentified automated traffic. Not every bot announces itself. Some scrapers and scripts spoof a browser user agent specifically to avoid being filtered. This traffic tends to show up as sessions with unusual patterns: zero time on page, no scroll, a single pageview with no referrer, or a burst of near-identical requests in a short window.
Human sessions with real engagement signals. Time on page, scroll depth, multiple pages per session, and return visits are far harder for basic bots to fake convincingly than a raw pageview count is.
Referral source. A human arriving from a search result, a newsletter link, or a social post behaves differently in aggregate than an automated fetch with no referrer at all. Traffic with no referrer and no engagement is a pattern worth treating with suspicion, not automatically trusting.
The practical takeaway: stop treating "total visits" as your headline metric. It's the least reliable number on the page. Any analytics tool worth using on a small blog should let you filter or segment by these categories, not just hand you one aggregate count and call it done.
Reading real engagement vs. crawl noise (subscribers and opens as the trustworthy signal)
If pageviews are noisy, what should you actually trust? The signals that are hardest for automated traffic to fake, because they require a deliberate action from a person, not just a fetch request:
Newsletter subscriptions. Someone typing in an email address and confirming it is a real human decision. Bots don't subscribe to your newsletter.
Email opens and clicks. These require a mail client belonging to an actual person opening an actual inbox. It's not a perfect signal (some opens are triggered by email client image prefetching, not necessarily the recipient), but it's a far cleaner read on genuine interest than a pageview count.
Return visits from the same person over time. A one-time crawl doesn't come back next week to read your next post. A subscriber pattern does.
Comments, replies, and shares that reference specific content. These take effort a script has no reason to spend.
This is the practical argument for building your audience around subscribers rather than around raw traffic. A pageview count tells you a URL got fetched. A subscriber list and an open rate tell you people are choosing to keep hearing from you, which is a much better proxy for whether the writing is working. If you're publishing on Floggy, the newsletter and analytics are part of the same Pro setup for exactly this reason: pageviews alone aren't the metric worth optimizing for.
Writing for readers vs. writing for crawlers: does anything change
Short answer: not much, and the instinct to change your writing because "AI reads it now too" usually points in the wrong direction.
The content that reads well to a human, answers a real question clearly, and is structured so the point isn't buried, is also the content that AI systems extract and cite most easily. Clear headings, a direct answer near the top of a section, concrete specifics instead of vague filler: all of that helps a human skimming the page and it helps an AI system trying to pull a self-contained passage out of your post. Writing for readers and writing for extractability point at the same output for the vast majority of posts.
What doesn't work is writing differently for the two audiences: stuffing extra keyword phrases in because you think a crawler wants them, or padding a post to hit a word count because you assume length signals authority. Neither search engines nor AI systems reward that, and it makes the post worse for the person actually reading it, which is the one audience whose opinion you can verify through subscriber growth and replies.
The one adjustment worth making deliberately: write passages that stand on their own. A paragraph that only makes sense in the context of three paragraphs before it is harder for a human to skim and harder for an AI system to extract cleanly as an answer. Front-load the point, then explain it. That's good writing advice independent of bots, it's just become more visible now that a wider range of automated systems are reading your posts too.
What to check in your own analytics this month
You don't need a research report to act on this. Open your own analytics and do the following:
Look at whether your traffic tool separates bot and crawler traffic from human sessions by default. If it doesn't, or if you can't tell which number you're looking at, that's the first thing to fix. A total that mixes both isn't a metric, it's a guess.
Check for sessions with zero engagement. Single pageview, no scroll, no time on page, no referrer. If a meaningful chunk of your "traffic" looks like this, treat your raw visit count with real skepticism before drawing conclusions from it.
Compare your subscriber growth rate to your pageview growth rate over the last few months. If pageviews are climbing but subscribers are flat, that's worth investigating, not celebrating. It may mean the visits aren't from people who intend to come back.
Look at your newsletter open rate as a health check, not just a vanity number. A steady or growing open rate among an honest subscriber base is a better read on whether your writing is landing than any traffic spike.
Revisit which posts you consider "successful" based on old pageview numbers. If those numbers were never separated from bot traffic, some of your past conclusions about what worked may be built on noise.
None of this requires guessing at industry-wide bot percentages you can't verify. It requires looking at your own numbers with the right skepticism and using the signals that actually require a human on the other end: subscriptions, opens, returns, and replies.
If you're setting up analytics for a new or growing blog, build this separation in from day one rather than retrofitting it later. Floggy's built-in analytics ships with your blog and newsletter on the Pro plan, so subscriber and open data sit next to your traffic numbers instead of living in a separate tool you have to reconcile by hand.


