Headless Browser Detection and Browser Fingerprinting: How a Filter Tells a Program From a Person

A headless browser runs JavaScript and renders pages like a real one, so simple checks miss it. Headless browser detection relies on the browser fingerprint and on automation traces that are hard to hide completely.

Bots and Fraud10 min read
Headless Browser Detection and Browser Fingerprinting: How a Filter Tells a Program From a Person
Contents
  1. What is a headless browser
  2. What is a browser fingerprint
  3. Signs of headless browsers and automation
  4. Antidetect browsers: where the line is
  5. False positives: in-app browsers and WebView
  6. How the browser check works in ArtisanClo
  7. How to build protection against headless browsers: practical steps
  8. Browser fingerprinting and privacy: what you should know
  9. Example: breaking down a suspicious visit
  10. Bottom line

A simple bot gives itself away by not running JavaScript. But ads stopped being checked by curl scripts long ago; they are checked by real browsers driven by programs. A headless browser like that loads the page, runs the code, waits, clicks — and to naive protection it looks exactly like a visitor. Headless browser detection relies on the browser fingerprint and on the automation traces the program leaves in its environment.

What is a headless browser

Headless means without a head — without a graphical window. It is the same Chrome or Firefox engine, just running in the background under a program's control. Popular automation tools such as Puppeteer, Playwright and Selenium can open pages, fill in forms, take screenshots and walk through a chain of redirects.

Who points headless browsers at your landing page:

  • ad platforms — they check the landing page the way a user will see it;
  • spy services — they save landing pages in full (more in the article on protection from spy services);
  • fraudulent traffic — it fakes clicks and even conversions;
  • competitor scrapers — they collect copy, prices and structure.

The core problem for a filter: a headless browser passes the JavaScript check. So subtler signals are needed.

What is a browser fingerprint

A browser fingerprint is the combination of characteristics a browser reveals to a site. Some are sent automatically, others are measured by a script.

Component What it is What can give it away
User agent string with the browser and OS name and version a template or outdated version
Language and locale the browser's language list a language that does not match the GEO
Time zone the system time zone a zone that does not match the IP country
Screen resolution, color depth, pixel density round server-style dimensions
Graphics graphics card, renderer, rendering quirks a software renderer instead of a real GPU
Fonts the set of installed fonts a bare server system
Hardware CPU cores, memory, touch screen desktop specs on a phone
API support which browser interfaces are available missing features a real browser of that version has

Fingerprints are used in two ways. For recognition — to tell that different visits came from the same device. And for authenticity — to tell whether this set of traits looks like a real device. For bot protection the second matters more; any anti-bot for a website is built on it.

Signs of headless browsers and automation

The webdriver flag

By the standard, a browser controlled by a program through an automation interface sets the navigator.webdriver property. It is the best-known signal and the easiest to hide: one line of code spoofs it. That is why serious protection never relies on it alone.

Traces of control

Automation tools leave their own objects, variables and telltale changes in how certain functions behave. There are many of them, they change with every tool release, and hiding all of them at once is hard.

Mismatch with what is claimed

A headless browser often betrays itself through contradictions:

  • the user agent says «Chrome on Windows», but the API set looks like a headless build on Linux;
  • it claims to be an iPhone but supports interfaces Safari does not have;
  • the screen is mobile-sized, but there is no touch input;
  • the graphics card shows up as a software renderer typical of servers.

A server environment

A minimal font set, no media devices, a round screen resolution and unnatural hardware values all point to a browser running on a server with no real user.

Timing and behavior

A program acts either instantly or with identical pauses. There are no mouse movements, no taps, no scrolling — or they are perfectly smooth. Faked time on page can be spotted too: what the script reports does not match the real timings.

Tip. No single signal on this list proves a bot on its own. The combination does: three or four independent contradictions in one visit almost never occur in real browsers.

Antidetect browsers: where the line is

An antidetect browser spoofs the fingerprint so that every profile looks like a separate real device. In affiliate marketing it is an everyday work tool: media buyers keep ad accounts in it so the platform does not link them together.

A traffic filter has to tell two cases apart:

  • an antidetect browser driven by a person — to the landing page it is essentially a normal visitor: a human moves the mouse, reads and clicks;
  • an antidetect browser driven by a program — the same automation, just better disguised; inconsistent spoofing and behavior give it away.

In other words, the fingerprint is not the only line of defense. A well-disguised bot is betrayed by what it does, not by how it looks.

False positives: in-app browsers and WebView

The most common mistake in fingerprint-based filtering is blocking in-app browsers. Facebook, Instagram and TikTok open links in a built-in browser based on WebView. It has:

  • an unusual user agent that includes the app name;
  • a reduced feature set compared to a full browser;
  • traits that overlap with automation traces;
  • often a lost referrer.

Yet most real social media traffic arrives through exactly these in-app browsers. So an automation signal in WebView must not be treated as a clear bot — only as one signal in the overall score. How the platform's click ID helps tell such a visit from a bot is covered in the article on fbclid, gclid and ttclid.

How the browser check works in ArtisanClo

In ArtisanClo the browser check is a separate step after the network checks. A script runs in the visitor's browser and reports what it saw; the page stays hidden until the decision arrives. Several protection toggles are responsible for this:

  • Block headless — catches bots and auto-clickers pretending to be a browser. With clear traces of automation it cuts the visit immediately, and the log shows the reason «Headless browser detected».
  • Require JS — a visitor without JavaScript gets a heavy penalty and is sent to a check. Simple bots and scrapers trip over this.
  • Live interaction check — looks at how the mouse or finger moves: people behave naturally, bots do not. It adds a short wait, so it is turned on for strict platforms.
  • Min time on page — a visitor who stayed less than the set time is sent to a waiting page and checked again.

In-app browsers of Facebook, Instagram and TikTok and WebView inside apps are recognized separately: an automation signal does not cut such a visit immediately but is weighed in the trust score along with the network, timing and referrer. For campaigns where almost all traffic comes from an in-app browser, the «Balanced» strictness is recommended.

Browser robots with proven signs and faked time on page go onto the shared bot list on the first catch: the next visit from that address to any flow of any customer gets the White Page right away.

In the statistics, rejections at this step are grouped into the «Browser check» share, and visits that did not run the script and did not come back go into «Check never came back». If it is the «Browser check» share that grows, that is a cue to review the strictness: this is the step where the filter errs most often.

The browser check works with both the JS tag and the PHP file connection. The difference between the two is explained in the article on how to connect a cloaker to your site, and the full list of checks is on the features page.

How to build protection against headless browsers: practical steps

  1. Start with the network. Most headless bots run on servers — blocking data centers cuts them before the browser check. See VPNs, proxies and data center IPs.
  2. Require JavaScript where it is safe. If the platform passes a click ID and traffic comes from regular browsers, requiring JS cuts scrapers without losses.
  3. Turn on headless blocking on all flows.
  4. Keep the live interaction check for expensive traffic and sources with a lot of advanced automation.
  5. Do not tighten rules for in-app browsers. Do not add extra harsh rules for social media traffic.
  6. Read the reasons. If «Headless browser detected» fires en masse on one platform's mobile traffic, that is a reason to investigate, not to celebrate.

The overall logic of layered protection is described in the article on how to filter bot traffic, and the list of typical tells is in signs of bot traffic.

Browser fingerprinting and privacy: what you should know

The word fingerprint is often associated with tracking: ad systems use fingerprints to recognize users without cookies. Anti-fraud has a different job. The filter does not need to know who the visitor is — it needs to know whether this is a real browser and whether it behaves like a person. That is why bot protection cares less about unique identifiers and more about whether the characteristics are consistent with each other.

This has a practical consequence. Browsers are gradually limiting what a script can measure: they add noise to rendering, unify user agent strings and hide hardware details. That is a problem for recognizing users, but less of one for authenticity checks: contradictions between what is claimed and what is real do not go away. Protection built on one or two magic signals, however, stops working over time — another argument for scoring on the sum of signals.

Example: breaking down a suspicious visit

Say the log has a visit with a reason related to browser automation. The click card shows:

  • network — cloud hosting;
  • the browser claims to be a fresh Chrome, and the device is a desktop;
  • there is no ad click ID;
  • the browser check came back with signs of program control.

Everything lines up here: a server network, no ad click, automation traces. This is a classic review robot or spy tool.

Now another visit: a mobile carrier, a social network's in-app browser, a click ID present, a weak automation signal. You cannot block this visit on one signal — most likely it is an ordinary person who opened an ad inside the app.

The difference between these two visits is the essence of scoring on combined signals: in the first, several independent signals contradict each other at once; in the second, just one, and even that is explained by in-app browser quirks. If the log has many visits like the second one and they end up on the White Page, lower the strictness for that platform.

Bottom line

A headless browser passes simple checks but leaves traces: automation flags and objects, fingerprint contradictions, a server environment, unnatural behavior. The browser fingerprint helps tell whether a visit looks like a real device, and behavioral checks whether it acts like a person. The main rule is the same as everywhere in anti-fraud: the combination of signals decides, not a single one — otherwise you lose social media in-app traffic along with the bots.

Frequently asked questions

01

What is a headless browser?

It is a regular browser engine that runs without a window and is controlled by a program. It loads pages, runs JavaScript, clicks buttons and saves the result. It is used for website testing, scraping, collecting ads and faking traffic.

02

What is browser fingerprinting in simple terms?

It is the set of characteristics a browser reports to a site or that a script can measure: version, language, time zone, screen, graphics card, fonts, rendering quirks. Together they form a nearly unique profile that can identify a device and show whether the browser is genuine.

03

What does navigator.webdriver mean?

It is a browser property that, by the standard, becomes true when a program controls the browser through an automation interface. It is the best-known bot signal and also the easiest to spoof, so serious checks never rely on it alone.

04

Can bot protection detect antidetect browsers?

An antidetect browser spoofs the fingerprint to look like a real device, and when a person drives it, it behaves like a person. To a filter such a visit is often indistinguishable from a normal visitor. It gets caught when the spoofing is inconsistent or when automation rather than a human is behind it.

05

Why does the Instagram or TikTok in-app browser sometimes look like a bot?

In-app browsers are built on WebView and differ from a full browser: some capabilities are cut and the user agent is unusual. Some of their traits overlap with automation signals, so such visits must never be blocked on a single signal, only on the combination.

Read next

See your traffic for real

Connect ArtisanClo to your site, see who actually arrives from your ads, and why every click got its decision.