User Agent Filtering: What a UA Is and Why It Is Not Enough Against Bots

User agent filtering is the oldest way to cut bots: the server reads the string the browser introduces itself with and decides based on it. Here is what that string contains, where you can trust it and why a UA alone is not enough to protect paid traffic.

Bots and Fraud11 min read
User Agent Filtering: What a UA Is and Why It Is Not Enough Against Bots
Contents
  1. What a user agent is and what it says
  2. How user agent filtering works
  3. Bot user agents: who introduces themselves honestly
  4. User agent spoofing: why a UA-only filter is weak
  5. Where user agent filtering is still useful
  6. Combining signals: the UA as one voice among many
  7. How it works in ArtisanClo
  8. Common user agent filtering mistakes
  9. Summary

A user agent (UA) is the string that a browser or any other program uses to introduce itself to the server in the header of every request: browser name and version, engine, operating system and sometimes device model. User agent filtering means the server reads this string and uses it to decide whom to let in, whom to show a different page and whom to cut as a bot.

The method is simple and cheap, so it still shows up in homemade scripts and "anti-bot" plugins. But the string is written by the client itself, and faking it is easier than faking any other signal of a visit. Below: what exactly is in a UA, where the string can be trusted, where it cannot and how to use it alongside other signals.

What a user agent is and what it says

A typical mobile Chrome string on Android looks like this:

Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Mobile Safari/537.36

It is easier to read in parts:

Fragment What it means
Mozilla/5.0 A historical compatibility prefix; almost every browser has it and it says nothing
(Linux; Android 10; K) Platform and OS; in modern browsers the device model is often replaced with a placeholder
AppleWebKit/537.36 (KHTML, like Gecko) The rendering engine and another nod to compatibility
Chrome/130.0.0.0 Browser and major version; the minor numbers have long been zeroed
Mobile Safari/537.36 A mobile layout marker

From this string a server usually extracts four things: device type (smartphone, tablet, computer, TV), operating system, browser and a sign that it is not a browser at all but a program or robot.

Note that the string contains almost no unique information. Millions of phones send an identical UA; that is exactly why browsers "froze" it, to make tracking people by it harder.

Client Hints: where the details are moving

Chromium-based browsers are gradually shortening the classic string and providing details through Client Hints, a set of headers including Sec-CH-UA, Sec-CH-UA-Mobile and Sec-CH-UA-Platform. Some arrive with every request, others only if the server explicitly asks. For filtering this helps in two ways: hints are harder to invent convincingly from scratch, and they can be cross-checked against the main string. A browser that calls itself Chrome on Windows in the UA but reports a different platform in the hints is behaving oddly.

There is a limitation too: Safari and Firefox support hints differently or not at all, so the absence of Client Hints says nothing on its own.

How user agent filtering works

In its simplest form it is a list of substrings: if the UA contains one of the banned words, the visitor goes to a stub page. A step up is a parser that splits the string into browser, OS and device, with rules on top of the result: "smartphones only", "no Smart TVs", "Android only".

It is important to separate two different tasks here:

  1. Targeting: selecting the target audience. Does this offer need desktop, does iOS fit? The UA string works fine here, because ordinary people do not fake it and a mistake is cheap.
  2. Bot protection: weeding out automation. Here a UA is almost useless as a standalone signal: anyone writing a bot starts by inserting a real browser's string.

Audience filters by device, OS and browser are covered in more detail in the article on geo and device filtering. The rest of this article is about the second task.

Bot user agents: who introduces themselves honestly

Some automated traffic really does introduce itself as it is. There are two different classes.

Crude scripts and tools. HTTP libraries, command-line downloaders and scrapers send their own UA by default, with the library name and version, or no UA at all. An empty UA string or a library name instead of a browser is a reliable sign: a real person from an ad does not arrive like that. Such visits are cut immediately and with almost no risk.

Search crawlers and "polite" robots. Search engines, link preview services and uptime monitors usually state their name and a link to a description honestly. For a regular site these are useful guests and are not blocked. But there is a caveat: scrapers often copy a well-known search crawler's string to get let in wherever a search engine is let in. That is why search engines themselves advise verifying a crawler not by its UA but with a reverse DNS lookup of the IP address: a genuine crawler comes from its company's network.

The takeaway: a UA filter is good at catching those who are not trying to hide. That is a noticeable share of junk requests, and cutting it is cheap. But it is not the part of bot traffic that eats your ad budget.

User agent spoofing: why a UA-only filter is weak

Any program that disguises itself even a little sends the string of a recent Chrome or Safari. That takes one line of code, or a toggle in a browser's developer tools. Three problems follow.

A fake is invisible in the string itself. A correct UA from a real browser and the same UA in a script's request look identical byte for byte. You cannot check the string against itself, only against other signals.

Lists go stale. A substring blacklist needs constant updating: new tools appear, old ones change names, and real browsers change the string format. A list that has not been updated for six months mostly catches things that stopped arriving long ago.

False positives on real people. In-app browsers, WebView, old phones and corporate browser builds send unusual strings. A rule like "anything that does not look like standard Chrome is suspicious" cuts exactly the mobile social traffic you pay for. In-app browser quirks are covered in detail in the piece on in-app traffic fraud.

Tip. If a UA rule fires on a noticeable share of paid traffic, it is most likely cutting real people, not bots. The bots worth building protection against arrive with a perfect string.

Where user agent filtering is still useful

Despite all of the above, the UA is a normal and necessary signal once you know its place.

  • Weeding out crude automation. An empty UA or the name of a library, command-line tool or scraping framework is an immediate "no". A cheap layer that takes load off the other checks.
  • Detecting device type for targeting and routing. Smartphone or computer, Android or iOS: for distributing traffic between offers, the UA string is usually enough.
  • Recognizing in-app browsers. The UA shows that the visit came from a browser inside an app. This matters not for blocking but the opposite: so as not to penalize such a visit for signs that are often false in WebView.
  • Cross-checking with other signals. The UA says "iPhone", but the screen, font set and script behavior say otherwise. The mismatch itself is the signal.
  • Analytics. Breakdowns by browser and OS in reports show where conversion sags and where a suspicious spike comes from.

Combining signals: the UA as one voice among many

Good protection does not ask "what UA does this visit have", it asks "does the UA agree with everything else". These are the signals it is usually checked against:

Signal What the UA is checked against
IP address and network A mobile browser from a hosting network deserves a closer look; more in VPN, proxy and data center IP detection
Request headers A real browser's set and order of headers differs from a library's, even if the UA string is the same
Client Hints The platform and browser brand in the hints must match the string
JavaScript execution A browser that claims to be a modern Chrome but did not run the script is not behaving like a browser
Browser properties in the script Platform, screen, signs of automation; more in headless browser detection and fingerprinting
Behavior Time on page, mouse movement and touches

Each signal on its own makes mistakes. Together they produce an assessment that makes mistakes far less often. This is how any modern smart traffic filter works: a single signal adds a little suspicion, and a confident conclusion comes from several of them agreeing. We covered the overall layered setup in the article on how to filter bot traffic.

How it works in ArtisanClo

In ArtisanClo the user agent string is one of the signals, not a standalone "list" rule. The decision on a visit is built from several steps, and the UA plays a part in each in its own way.

Network and request. The first step looks at the address, network, provider, headers, click ID and flow rules, before any browser check. Programs instead of browsers and platform robots that build link previews are cut here: such visits always get the White Page, whatever the settings, and the log shows the reason "A program, not a browser" or "Search or platform crawler". This step affects real people the least.

Audiences. The flow's filtering settings have Devices, Operating systems and Browsers lists with allow and block options. Among devices, Bot / crawler and Unknown are listed separately. This is targeting: a visitor who does not pass the list sees the White Page with a clear reason.

Browser check. Anyone who sent a real browser's string goes through a script check, and its response shows whether the visit looks like automation. Blocking headless browsers cuts visits with clear traces of automation straight away; less obvious signs go into the trust score together with the network, VPN, the presence of a provider and a referring site.

In-app browsers. The in-app browsers of Facebook, Instagram, TikTok and WebView in apps are recognized separately. A sign of automation does not cut such a visit outright; it is weighed in the trust score together with the network, timing and referrer, because in these browsers it is often false.

Transparency. Every click in the log has a decision reason from dozens of possible ones, and the Visit report link shows all signals and conclusions for a single visit. If you suspect the filter got it wrong, start with why a click went to the White Page. All filtering features are on the features page.

Common user agent filtering mistakes

  1. Building protection on a substring blacklist. It only catches those who do not hide and needs endless updating.
  2. Blocking everything "non-standard". Social in-app browsers and old devices get hit: real paid traffic.
  3. Trusting a string that looks right. A perfect modern browser UA is exactly what a well-made bot sends.
  4. Confusing targeting with protection. "Android only" is a decision about the offer's audience, not a way to get rid of bots.
  5. Not looking at filtering reasons. If you cannot see which rule fired, you cannot tell whether the filter is cutting bots or people. The signs that make bot traffic visible in reports are collected in bot traffic signs.

Summary

The user agent string is what the client says about itself. You can trust it when the stakes are low: device detection, traffic routing, analytics. You cannot trust it as proof that you are dealing with a person. User agent filtering is good at removing crude scripts and empty requests, but real protection starts where the UA is cross-checked against the network, headers, Client Hints, script execution and behavior. Then a faked string does not help a bot, and an unusual string does not hurt a real visitor.

Frequently asked questions

01

What is a user agent in simple terms?

It is a short text business card that a browser or program sends to the server with every request. It says what the program is, which version and which system it runs on. The server uses it to choose the layout, collect statistics and sometimes decide whether to let the visitor in.

02

Can a user agent be faked?

Yes, very easily. In any HTTP library it is one line of configuration; in a browser it is a developer tools option or an extension. So a UA matching regular Chrome proves nothing, while an obvious mismatch between the UA and the visit's other signals is a strong signal.

03

How do I find my user agent?

The easiest way is to open the browser's developer console and run navigator.userAgent, which prints the full string. You can also see it in the Network tab in the headers of any request. Many checker sites show it too, but the browser's built-in tools are enough.

04

Why aren't search crawlers blocked by user agent?

For a regular site search crawlers are useful: without them the page will not appear in search results. Besides, their UA is often faked, so search engines themselves recommend verifying a crawler not by the string but with a reverse DNS lookup of its IP address. Letting them in or not is the site owner's decision, not the job of an ad traffic filter.

05

Will Client Hints replace the user agent string?

Partly. Chromium-based browsers already shorten the classic string and move details into Sec-CH-UA headers, and some of them reach the server only on request. Other browsers support them only partially, so in practice sites read both sources and cross-check them.

Read next

See your traffic for real

Connect ArtisanClo to your site, see who actually arrives from your ads, and why every click got its decision.