Smart Traffic Filter: Inside a Bot Detection System

A smart traffic filter does not just check a visit against a list of rules. It weighs dozens of signals, learns from history and adapts to your traffic. Here is what is behind it, where it gets things wrong and how to keep its decisions under control.

Bots and Fraud11 min read
Smart Traffic Filter: Inside a Bot Detection System
Contents
  1. How a smart traffic filter differs from a rule list
  2. What a bot detection system is made of
  3. Learning from history: how the filter gets smarter
  4. A shared list of proven bots
  5. Automatic filter tuning
  6. Risks: when a smart filter blocks real people
  7. Transparent decisions: without them a smart filter cannot be trusted
  8. How the smart traffic filter works in ArtisanClo
  9. Summary

A smart traffic filter is a bot detection system that decides on a visit not with a single rule but with a combination of many signals: network, browser, behavior and history. It weighs those signals into an overall trust score, learns from past data (who turned out to be a bot and who brought money) and tunes protection to your specific traffic.

A simple filter answers the question "was a rule broken?". A smart one answers "how much does this visit look like the person we actually want?". Below we look at what such a system is made of, why it beats a plain rule list, where its weak spots are and why you should not trust it without transparent decisions.

How a smart traffic filter differs from a rule list

A classic filter is a list of conditions: country not on the list, show the stub; IP from a data center, stub; no JavaScript, stub. That works for the obvious cases and breaks down on borderline ones.

Rule list Smart traffic filter
How it decides One triggered rule means rejection Signals add up to a trust score
Borderline visits Either blocks everyone or lets everyone through A doubtful signal adds suspicion but does not decide alone
Adaptation Manual edits only Learns from history and shared experience
New bots Get through until someone writes a rule Similarity to known bots catches new variants too
Risk Predictable but crude Finer, but needs control and explanations

Take a visit through a VPN. A rule list has to choose: block every VPN and lose some real people, or let every VPN through and admit bots. A smart filter treats the VPN as one signal among many. If everything else looks right (a real browser, the script ran, the platform's click ID arrived), the visit passes. If the VPN comes together with a hosting network and traces of automation, the accumulated suspicion settles it.

Note that a smart filter does not cancel rules. Clear-cut cases such as a program instead of a browser, an address from a blacklist or the wrong country are still cut immediately. The "smart" part is for the middle ground, where a rule errs in both directions. We covered the overall layered setup in the article on how to filter bot traffic.

What a bot detection system is made of

Any modern bot detection system has four parts.

1. Signals

The more independent signals there are, the harder it is for a bot to fake all of them at once:

  • Network: address type (residential, mobile, hosting), VPN and proxy, provider, ASN; more in the article on VPN, proxy and data center IP detection;
  • Request: headers, the user agent string, the platform's click ID, the referring site; why the UA string alone proves little is covered in the piece on user agent filtering;
  • Browser: JavaScript execution, signs of automation, consistency of properties; see the article on headless browser detection and fingerprinting;
  • Behavior: time on page, mouse movement and touches;
  • History: whether this address has been seen before and how earlier visits ended.

2. Trust score

Signals are reduced to a single value, the trust score. Every suspicious signal lowers trust, and if too many suspicions pile up, the visit is filtered out. The weight of each signal depends on context: a missing referrer is suspicious for some platforms and perfectly normal for others.

3. Learning

Learning answers a question rules never ask: what do your bots and your buyers look like? The model looks at history (which visits ended in a payment and which turned out to be bots) and finds patterns a person would never write down as a rule.

4. Memory

A bot that was confidently caught once will most likely come back. A shared list of proven bots means it does not have to be checked again: the next visit from the same address is cut immediately. The more traffic the system sees, the more useful this memory becomes, which is why it is shared across all of a service's customers.

Learning from history: how the filter gets smarter

Learning is the strongest and the trickiest part of a smart filter.

It is strong because the model sees what the rules missed. A bot that perfectly imitates a browser still does not behave like your buyers: different hours, different networks, a different combination of signals. A model trained on the account's history notices that mismatch.

It is tricky because the model learns from what it is shown. If the history has few conversions, the model has a vague idea of what a buyer looks like. If the traffic changes (a new platform, a new geo, a new creative), old patterns stop working. That is why good learning systems have three safeguards:

  1. A shared model to start with. While your own data is thin, a model trained on many accounts does the work.
  2. Regular retraining. The model is updated on fresh data instead of living for years on an old snapshot.
  3. A conversion fuse. If the model starts cutting a noticeable share of paying visitors, it switches off instead of continuing to cut revenue.

A shared list of proven bots

A bot list works like herd immunity: whatever was caught at one customer no longer gets through anywhere. It carries an obvious risk, though: adding a real person to it. A mobile carrier address that a bot used today may be assigned to an ordinary subscriber tomorrow.

So it makes sense to put only proven bots on such a list, not everyone suspicious, and to keep entries for different periods: hosting and data center addresses rarely change hands, while residential and mobile ones change often. For more on how address lists work, see the article on the IP blacklist.

Automatic filter tuning

The last stage is a system that not only decides on each visit but also changes its own settings: it turns on protection that would catch bots on this traffic, or removes a rule that cuts paying visitors. That saves a media buyer time, but it needs three limits:

  • Small steps. One change at a time and no more often than a set interval, otherwise you cannot tell what affected the result.
  • Untouchable decisions. Geo, devices and schedule are business decisions, not protection, and automation must not touch them.
  • Undo. Every change must be visible with an explanation and reversible in one click.

Risks: when a smart filter blocks real people

The main risk of any filter is false positives. In a smart filter they are less visible, because the decision is made of many small contributions and does not come down to one clear rule.

Signs that the filter is cutting people:

  • conversions drop as filtering grows, even though the platform and creatives have not changed;
  • only a very small share of visits reaches the offer, which is a reason to check the settings, especially on paid traffic with a click ID;
  • filtering grows at the browser check step, which is where filters make mistakes most often;
  • the filtered traffic contains many social media in-app browsers and mobile networks.

There is no reverse rule, though: a high pass rate on its own does not mean weak protection. On a paid source with a verified click ID, most visits may reach the offer, and that is normal. How to tell good traffic from bad in reports is covered in the article on bot traffic signs.

Tip. Before tightening protection, look at whom it would filter out without actually filtering anyone. Observation mode costs little, and lost conversions never come back.

Transparent decisions: without them a smart filter cannot be trusted

The more complex the system, the more important it is that it explains every decision. Otherwise you cannot tell protection from a malfunction. The minimum level of transparency:

  1. A reason for every rejection. Not just "blocked" but something specific: hosting network, signs of automation, wrong country, resembles the account's bots.
  2. A breakdown of a single visit. Which signals arrived and what conclusions were drawn from them.
  3. The step at which traffic was filtered. Network, browser or behavior; this shows which layer is getting it wrong.
  4. A log of automatic changes. What changed on its own, when and why.

If a system cannot do at least this much, the word "smart" in its description is worth nothing. For more selection criteria, see how to choose a traffic filtering service.

How the smart traffic filter works in ArtisanClo

The ArtisanClo filter is built from the same parts, and each of them comes with a clear explanation in the dashboard.

Rules for the obvious. Some checks cut a visit immediately: IP blacklists, an obvious robot browser, ad review services, programs instead of browsers, audience rules for country, device and schedule.

A trust score for the borderline. The remaining signals (VPN, data center, IPv6, a missing provider or referrer, no JavaScript and others) add up to a trust score. Protection toggles decide how much each signal weighs, and the strictness levels Soft, Balanced and Strict are matched to the traffic source's platform.

Extra guard: learning from history. From the Professional plan, Diagnostics has an Extra guard tab. The model learns from your account's history: whom you were paid for and who turned out to be a bot. It only checks visits the other rules have already let through, and sends visits that resemble your bots to the White Page. While your own data is thin, a shared model does the work. The model is retrained every night and switches itself off if it starts cutting a noticeable share of conversions.

The shared bot list. Ad review services, programs instead of browsers, robot browsers with proven signs and platform preview robots go onto the shared list the first time they are seen, and their next visit to any flow of any customer gets the White Page straight away with the reason Known bot address. Real visitors with a questionable network, a VPN or no JavaScript are never added. Data center and hosting addresses are kept practically indefinitely, residential and mobile ones for a limited time. In addition, an address caught on automation at several customers at once goes onto the shared IP blacklist within an hour.

Auto-tuning: automatic adjustment. From Professional, Auto-tuning looks at the flow once an hour and makes no more than one change at a time: it turns on protection if it would have caught enough bots without touching a single payment, or removes a rule that was cutting visitors with payments. It never touches geo, devices or schedule. Every change appears in the auto-tuning log with an explanation and can be undone, and after an undo Auto-tuning leaves that setting alone.

Transparency. Every click in the log has a decision and a reason from dozens of possible ones. Statistics show the step at which traffic was filtered: Network and request, Browser check or Check never came back. The Visit report shows all signals and the final decision for a single visit, while recommendations in Diagnostics replay past visits through the current settings and show which filter is cutting traffic with conversions. Shadow mode on the PHP file and in combination with Keitaro or Binom shows whom the filter would have cut without cutting anyone. If a click went the wrong way, start with the article on why a click went to the White Page.

All protection features are collected on the features page, and plan terms are on the pricing page.

Summary

A smart traffic filter is rules for the obvious, a trust score for the borderline, learning for what rules cannot describe and a shared memory of proven bots. Its strength is that it sees combinations of signals and adapts to your traffic. Its risk is that its mistakes are subtler and harder to notice. So choose a bot detection system not by promises to "catch every bot" but by how clearly it explains each decision, whether it can observe without filtering and whether it stops on its own when it starts cutting paying visitors.

Frequently asked questions

01

What is a smart traffic filter?

It is a bot detection system that decides the fate of a visit not by one rule but by a combination of signals: network, browser, behavior and history. It can learn from past decisions and adapt to specific traffic. The main difference from a simple filter is that it handles borderline cases, not only obvious ones.

02

Can a smart traffic filter block real people?

Yes, like any filter. The more a system relies on probabilistic conclusions, the more closely you need to watch false positives: check which reasons filtered out traffic and compare that with conversions. A good system notices on its own when it starts cutting paying visitors and rolls the change back.

03

Do I need to configure a smart bot filter manually?

The basic decisions are still yours: which countries and devices you want, how strict filtering should be and which platform the traffic comes from. Learning and automatic tuning refine protection on top of those settings, but they do not replace knowing who you want to see on the offer.

04

How much data does a bot detection model need to start learning?

It depends on the system, but it always needs history that contains both bots and confirmed conversions. While your own data is thin, sensible systems run on a shared model trained on traffic from many accounts and switch to a personal one once enough examples accumulate.

05

Why should a bot detection system explain its decisions?

Because without an explanation you cannot tell protection from a malfunction. If you see that a visit was filtered for, say, a hosting network or signs of automation, you can check the decision and relax a specific rule if needed. If all you see is blocked, you either trust the system blindly or switch it off entirely.

Read next

See your traffic for real

Connect ArtisanClo to your site, see who actually arrives from your ads, and why every click got its decision.