Not All That Flows Freely Is Clean
Let’s set the dramatic scene…
You’ve got all the serious players in the room: your best analysts, your General Counsel plus a top tech M&A attorney, your ace dealmaker, the works-or maybe you are one of these people at a big investment fund, settling into a big chair at this sit down.
It wasn’t even hard to Tetris this eye-wateringly expensive meeting onto all these usually impossible calendars, because what’s being exhibited is a unicorn among unicorns, and people made the time.
There’s a digital ad network, or maybe a martech platform, that’s immensely profitable and growing rapidly. They’ve positioned themselves around results, and they’re delivering, clearly well enough to be judged a real driver of long term growth the way their existing client revenue growth is popping.
Google, Meta, TikTok…with names like those having made all their money in the space, the potential investment yield of of a digital performance marketing player isn’t a point that needs to be belabored.
The company’s reported financial figures tell a clear success story, but the job of a room like this is to look beyond what’s obvious to everyone else in this present moment. You’re seasoned pros at this, so the key question comes out almost immediately:
What’s their moat?
How do you know this success isn’t temporary, or something that could be grabbed by a fast follower or two to turn this nice little blue lagoon you’ve spied into roiling red puddle of blood?
That’s a great question, but identifying that moat shouldn’t be the end of your diligence journey.
Your prospective acquisition target may not be because it’s a majestic castle surrounded by a sparkling blue wall of water, and instead be a dark compound of degeneracy and villainy sitting in the middle of a foul swamp it’s created by pumping out its own rotten sewage.
We’ll help you find out if the trench that makes barbarians sink is giving off an awful stink, but first, we feast…on recent relevant news!
Important Happenings & New Notions
Nielsen buys DoubleVerify
A longtime large player in the digital ad verification and brand safety space is now owned by the world’s most famous, if not most recently fortunate, advertising measurement company.
Roblox stock down 70% YoY
More than any other company in the world, Roblox lives at the intersection of the questions “how much internet use is done by literal children” and “is that bad?” Some predictable stock market headwinds have started to blow.
AI fuels mobile ad fraud fire
IAS released a report about The Papyrus Network, a group of e-reader apps that send fraudulent traffic to a vast network of synthetic content sites generated by LLMs, which have been a godsend for mass creation of online fraud fodder.
Prediction markets shop for trust
If you’ve read deeply about prediction markets and sports, you know there are deep concerns about the verifiability of propositions, leading to possible outright manipulation. Bettors and the companies themselves may have more redress with a better data and trust layer, which a deal with Genius Sports aims to deliver.
Tl;dr
Modern digital advertising and marketing tech benefit immensely from being able to persistently identify users; it is often times a desirable way to build a “moat” for the business, or to at least enhance it.
They face additional pressure to do so from companies like Google and Meta, that have a massive material advantage at this due to their large logged-in user bases and many web and app properties.
Parties without this advantage may need to be more reliant on probabilistic attribution-and unfortunately, this type of user identification techniques might lead to liabilities you should learn about in order to ensure you can guard your clients and investment portfolio against them.
The “Problem” Being “Solved”
Here’s a helpful hypothetical example:
Let’s pretend we are a mobile ad network that’s running a campaign for one of Netmarble’s Game of Thrones mobile games, currently in an optimal marketing window given that the House of the Dragon season finale has just aired.
The first thing we’re going to do is run CTV ads during any Game of Thrones show, that’s a no-brainer; for good measure, let’s throw in some ads during general fantasy content on CTV also.
The whole point is to sell the user a mobile app, and here’s where the dilemma begins-this would happen on their mobile phone and would be best accomplished with a mobile app acquisition ad that appears on said phone.
A standard ad network, using standard tracking, can’t determine which mobile devices belong to viewers of the TV show or related content; we’re going to be stuck running well-targeted TV ads that rely on mobile download QR codes, and mobile acquisition ads with only mobile centric targeting data, which likely leads us to high spending mobile gamers with no particular interest in the Game of Thrones IP.
If we know you watched a lot of Game of Thrones on your TV but can’t find your mobile device? Oof. Here’s the user-side view:

This is what our poor Game of Thrones viewing user sees-the standard specious lootbox mobile game inventory we blast out to mobile users we don’t know much about. Alas!
If we have observed your mobile device making a ton of high revenue mobile gaming purchases but can’t find your TV and the massive trove of low quality anime you watch?

You’re a high value mobile app money spender and so we’ve saved our best client’s flagship product for you-too bad you haven’t watched a western media property since 4th grade and we’re wasting an ad slot that would be better suited to some sort of abominable waifu collection machine!
Now, if we were a network able to successfully stitch together the identity and behavior of these users, we could properly bucket our premium Hollywood USA major IP mobile game ads, and our gacha slop, properly.

This is the kind of harmonized, cross-device ad experience that is still achievable at scale, ethically, only by Google and a few other companies, honestly, for reasons I will detail.
We could target users with interest based on their TV content viewing, and target them at the point of purchase on their phones, using highly effective ad formats native to the device and directly connected to the app store.
The inverse use case would also be possible for products that users demonstrate fitness to buy on mobile devices but that are advertised effectively on TV screens.
This trick is not so simple for platforms that are not a Google or a Meta or an Amazon, with lots of logged in devices and services allowing them to paint a complete picture of a user’s digital life deterministically.
If what you’re looking at acquiring or investing in is more of a “tier 2” type digital ad network, what you’re probably looking at is someone doing probabilistic attribution, which speaking probabilistically, is a probable problem.
Deterministic vs. Probabilistic Attribution
The way most ad platforms track users is deterministically, which is a word we’re using in a funny way here, because really we’re using it just to say it is NOT being done probabilistically.
This means that when a user “appears” to the network’s system in an ad call or a conversion pixel or otherwise, it’s just going to check for a basic parameter conveying a user’s identity. If my user session data has an email set to “[email protected],” or my phone number, or a unique identifier number generated by the ad network, than it knows I am me, complete with the context of my prior behavior it has seen, and that’s all there is to it.
However, if I have made no user identity data available to a deterministic tracking system, it will treat me as a new user-and this is where martech and ad networks face a major shortcoming if they’re using a purely deterministic system.
Probabilistic tracking is when, if faced with this uncertain identity, an ad network tries to determine if an “unknown” user is actually a “known” user by calculating the probability that this could be any number of users already known to the network.
Imagine you’re at a business networking event, and you see a circle of people, three of whom have nametags, and one who doesn’t.
However, the person with the name tag is about 6’3” tall, has short brown hair, black frame glasses, and is wearing red Puma sneakers, which you recall all being the traits of a person in your industry named “Jacob” you met at similar events previously.
It would be very reasonable to infer that this is “Jacob.”
Innocent enough! We all do this literally every time we see a person, technically, because in most of our everyday lives, people we encounter aren’t wearing name tags.
If you are married, when you wake up in the morning and look at your spouse, you are probabilistically attributing them as your life partner, unless you’re life partnered with a Pokémon who says their name first thing in the morning and hundreds of times throughout the day.
That situation actually sounds kind of convenient…maybe you should give Miss Cigarettes a call?
Describing probabilistic attribution this way is probably making it seem completely harmless and maybe even necessary, so let me give you a hypothetical from the dark side of the spectrum.
You’re at this same business networking event and walk up to a person you' haven’t really met, in any conventional sense-you’ve just seen them from across the room once or twice. As you approach them, you say:
“Hey Jacob, there’s a limited time event happening in the Game of Thrones Kingsroad mobile game offering a large amount of in-game currency with a standard battle pass purchase. You probably have a little extra cash now because your car loan was just fully paid off, but could use these savings soon given that your fertility testing app usage suggests you’ll have some serious diaper bills soon.'“
You’re able to do this because for several months you’ve been going through their garbage bin, and hired someone to hack their phone so you can see everything they do on it.
While this example is a little dramatic, it illustrates the alleged bear case against probabilistic attribution: that it’s a dangerously careless maximalist data vacuum that uses potentially problematic methods to consistently identify a user online, and along the way ties them to large volumes of low-quality or consent-questionable data that many privacy advocates argue should not be used to inform the ads they see.
While there are many approaches, often used in concert, to persistently identify users, the one at the center of the controversy due to its import and potential problems is device fingerprinting.
Dusting for Data
The device fingerprint, which is a body of information a machine generates that makes it persistently identifiable, has been used since before it even had a name in the 1990’s, and crystalized clearly as a concept to be examined in the 2010’s.
It’s become more effective as the methods have become increasingly sophisticated, principally through the collection of an ever-expanding list of parameters and even arguably innovative ways to generate them.
The best way to understand this, and I sincerely recommend you just do this right now to wrap your head around this is if you haven’t yet, is to check out your own fingerprint on the device you’re using right now.

Anyone with basic web development knowledge can do things like make a device render an image (“Canvas”) and look at how it was done, which combines so many device parameters that it does 99.98% of the work alone.
Clearly, effective device fingerprinting is not difficult. It’s the propriety of this practice in the eyes of a number of key parties that is a potential issue.
Who Takes Issue?
Well, let’s start with the big one that will probably surprise nobody: Apple.
We detailed the incredible importance of ads and ad tech working well on Apple’s devices and all of the challenges with that in a previous article.
Due to their privacy policies and the accompanying technology on the iPhone and within Safari, any platform getting a little too daring with their probabilistic attribution of users on these platforms might run afoul of arguably the world’s most pivotal mobile ecosystem player.
On the American legal side, the CPRA, the 2020 amendment to the landmark CCPA, includes unique identifiers and probabilistic identifiers as personal data, which may cover device fingerprinting in practice, but does not explicitly classify it as personal data collection.
It’s possible your potential investment destination is operationally and legally covered for any probabilistic attribution it does, but you will want to make sure you check all the right places to ensure they’re checking all the right boxes.
Where Probabilistic Problematicness Hides
The internet using public and researchers with the access to the same data as anyone else have seen their information layer come to represent a significantly thinner slice of user tracking over the past decade or so.
This is for many reasons: our online behavior has become more centered around apps and “walled gardens,” cookies have been degraded by privacy technology, ad blocker usage has reached levels around 30%, etc.
Another lesser known, but key, reason for this (that is also a reaction to some of the developments I just mentioned) is the rapid proliferation of server-side analytics, in large part to improve ad platform tracking fidelity, among other factors; one of the most well-known examples of this is Meta’s Conversions API, or CAPI for short.
This takes user tracking data out of browser-visible web cookies, and moves it into a slightly less transparent layer, away from cookies and web clients and into the domain between various organizations’ servers via API.
If you’re excited about an ad network’s ability to successfully target the right users to win big performance budgets, you’re going to want to audit their data that flows through these APIs, on ad calls and on any tracking infrastructure, and maybe also delve into their web pixel and app SDK implementations.
The good news is that when you’re considering what kind of framework to use to assess all this, a potentially suitable one exists, and it’s possible that whatever piece of martech or ad network you’re looking at already uses it: Apple’s Required Reason API.
This is basically what it sounds like: an API that requires structured data stating why a particular piece of information is required for collection by the party that’s collecting it, within which the only reasons one can list are deemed policy acceptable.
It’s worth stating here that it is possible that certain parties could purposefully misstate the real reason for collecting something, or give entirely false reasons to collect data they don’t need for any actual acceptable reason.
It’s a start, though, and you could do your own audit of the reasons given in the Apple API, and use the core idea as a framework for auditing all other data streams!
What Is A Positive Finding?
There are good reasons to collect user information for apps, and believe it or not, there are even non-advertising and marketing reasons to do device fingerprinting, even by ad tech and martech platforms.
Many of the same things that can be used for device fingerprinting can legitimately be collected for apps, including some owned by ad platforms, or part of ad networks they pass proprietary data to. Location based apps need telemetry data, and apps that deliver complicated experiences need complicated readouts on exactly how a device is handling many different factors and what its technical specs are.
Device fingerprinting, done explicitly, can also be used legitimately in antifraud efforts. A great deal of internet fraud is accomplished by hijacking devices, and/or spoofing one device to make it look like a different device entirely.
The best way to fight this is to make as accurate a graph of devices used for fraud as one possibly can, and to aggressively profile any and all devices on your ad network to see if they exhibit characteristics typical of a fraudulent device.
There’s a key question to ask here and a viable way to investigate the answer.
Much antifraud work is handled by companies that specialize in it, and often integrate with ad tech and martech platforms to handle this at both the general platform and individual client levels. It’s typical for a large platform to do one or all three of these things:
Have an agreement where antifraud services are provided to the whole platform for every impression and data event, for the benefit of the platform and their clients indirectly.
Have integrations that present, and arguably promote, client opt-in to these services, either gratis or for an additional fee that is passed along to them via their platform bill. There’s usually a tick box in account or campaign setup to enable this.
Allowing clients to bring their own solutions via minor engineering enablement. This may take the form of simply allowing tracking URLs and view through attribution pixels, or a flexible API to deliver data to different verification vendors, among other platform features.
If your prospective acquisition target is doing device fingerprinting or collecting extra data ostensibly for anti-fraud purposes, you should ask how effective this work is, and if it’s truly part of the ideal approach to fighting fraud on the platform.
It may be that the platform is building and maintaining a toolkit it makes more sense to use an industry standard specialized solution for.
Unfortunately, it’s also possible that if they claim that they’re using these solutions and collecting this data “just in case” or to check anti-fraud homework, it doesn’t actually reflect a genuine operational need, and they’re doing the kind of device fingerprinting that runs afoul of parties that could negatively impact their valuation if it were found out. This would require additional diligence to confirm.
Nobody wants to be the proud new owner of an app network that is about to catch a ban from the Apple App Store!
Get the accountable person to quantify what this fingerprinting is doing for them in terms of the fraud it’s catching and determine if it’s worth the risk.
Fare Thee Well, Acquirer
I leave you with this Olde Medieval Blessing penned by the Knights of the Cap Table:
May every moat you cross by boat be clear, clean data on which to float /
And if it stinks just deign to think, does the API send the kitchen sink?

