Fake Photos and Fraud
Friday, 24 August 2018
Pictures have always had power. But with the ubiquity and speed of social media, pictures have more influence on public opinion than ever before. It is very easy for a fake picture to go viral, and very difficult to disseminate any kind of correction.
Lots of people want a one-button solution. Something that will quickly tell them whether a picture is real or fake. Sadly, there is almost never a simple yes/no answer. But this desire for a fast, simplified solution opens the door for lots of snake-oil solutions and charlatan products.
This time, the product is called SurfSafe. This is a brand new Chrome plugin that claims to catch fake news by spotting fake photos. News outlets like Wired, BoingBoing, and Mac Observer recently touted this wonderful, new plugin. (I can only assume that none of those reporters actually reviewed the product before summarizing the PRweb press release.)
TL;DR: In my professional option, stay away from the SurfSafe plugin. It has serious privacy violations and does not do what it claims.
After that, you start surfing the web. Every picture gets a little icon above it. Clicking on the icon queries the SurfSafe servers for information about the picture. It returns the number of sightings, links to the sightings, and number of people who have reported the image. (Reports identify the number of people who think it is fake.)
Hovering the mouse over each picture also reveals a "Report" button. This is where you can report a picture as being "Propaganda", "Misleading", or "Photoshopped". This way, other people who see the same picture can quickly see how it has been reported.
Then I looked at my web browser. With Chrome, you can press Control-Shift-i and pull up the developer panel. Then you can click on the "Network" tab and see all of the network requests. Here's what my FotoForensics service normally looks like:

Normally, there are 5 network requests: the web page, style sheet, banner picture, fonts, and my little IPv4/IPv6 experiment.
And here's what happens when I have SurfSafe enabled:

Every web page gets passed through the SurfSafe extension. However, that final 'query' is the troubling issue. It submitted the URL for the banner image to SurfSafe. If my web page has 5 pictures, then all five image URLs would be sent from my browser to SurfSafe.
The number of privacy issues here is stunning:
For my own testing, I visited a couple of web pages on the public FotoForensics site. Then I queried SurfSafe about my banner. Here's the result:
My banner logo was found on three sources. Those are the three web pages at FotoForensics that I visited over HTTPS. (I checked my logs: SurfSafe only visited those pages after I went there with my browser.) Granted, on their public "sources" list, they removed the URL parameters from the source URLs (meaning that any URL requiring parameters will be a broken source if you click on it). However, they still list the web page and the page's HTML title; the title alone may give away too much personal information.
In any case, if you use the SurfSafe plugin, then you are feeding the URL of every picture you come across into their service for their collection. This is a huge privacy issue.
I uploaded the picture to FotoForensics. The metadata clearly states that the picture had been processed using Adobe Photoshop CS3. So, I used SurfSafe to 'report' the picture as 'photoshopped'. I reloaded the page and now there is one report for the picture (mine). However, nowhere does SurfSafe mention the cause of the report. There's just "Reports 1".
Clicking on Kavanaugh's Wikipedia picture takes you to Wikimedia Commons, where another copy of the picture is hosted. And when I say it's another copy, I mean it has the exact same SHA1 checksum -- bit-per-bit, it is the same picture. At Wikimedia Commons, SurfSafe reports 4 sources: 3 at Wikipedia (which appear to be the same URLs) and one at Wikimedia. However, the Wikimedia page doesn't have any "reports".
I then went back to the Wikipedia page and reloaded. Now it has 3 sources listed and one report. I waited about a minute and reloaded again: 4 sources and still one report. So even through a photoshopped picture is a photoshopped picture wherever it lives, SurfSafe only tags the photoshopped report to the one URL where I reported it as photoshopped.
What this means: SurfSafe won't flag a picture as being propaganda, misleading, or photoshopped unless someone else already flagged the picture at that specific URL. And since many web sites, including Facebook and Twitter, use user-specific URLs, it's very possible for other people to see the same picture and not be told that it's been altered. For SurfSafe to claim that they help detect fake photos is misleading at best.
misleading/photoshopped has a name: it's a mechanical turk. The mechanical turk will be just as accurate as the humans who categorize the photos. And the results will be just as consistent as the various human opinions. (Unfortunately, most humans are really bad at evaluating photos just by looking at them. And many people have different opinions.) Crowdsourcing the determination about whether a picture is real or fake (propaganda, misleading, photoshopped, etc.) is neither scientific nor accurate.
Crowdsourcing usually results in the most popular solutions. Consider the case of Boaty McBoatface. In March 2016, the Natural Environment Research Council ran a poll to name one of their ships. Through the miracle of crowdsourcing, the name "Boaty McBoatface" was selected. (This created a controversy when the NERC decided to assign the name to a different vehicle -- one with a less public image.) However, this shows how crowdsourcing works: it's not always the best solution.
There are other issues with this crowdsourcing solution. For example, nobody vetted me. I didn't have to login or create any kind of unique ID to submit my 'report' to SurfSafe. This means that anyone can poison pictures at SurfSafe. It doesn't take much for a bot to submit thousands of reports against lots of pictures. (And you can always use a list of open proxies if you want to come from ten thousand different addresses.) Organizations with specific agendas can easily have real pictures marked as untrusted in order to cause confusion or attempt to discredit negative publicity.
Perhaps SurfSafe should consider a Facebook-style user rating system. Facebook users who forward known-false stories will earn a lower trustworthiness score than people who forward vetted stories. This reputation management system is Facebook's attempt to combat fake news. The idea is that users with low trustworthiness scores won't be able to propagate new stories very far. If you have a reputation for promoting false stories, then any new content from you will be treated with a healthy amount of skepticism. Then again, SurfSafe will still need a way to vet stories before assigning user trustworthiness.
The problem is that no news outlet is perfect. Trusted news outlets may occasionally make mistakes. (Some news outlets have more mistakes than others. Right, Fox?) This is particularly important in this era of rapid news coverage; for many news outlets, speed is more important than accuracy. Without vetted news reporting, we end up with situations like the Boston Marathon bombing, where the wrong person was named by some media outlets.
There's also a false equivalency and false balance problems that some media outlets experience. A few journalists think "fair and balanced" means that they must provide equal coverage to both sides of a reported topic. If they say something negative about Nazis, then they must also find at least something positive to say. Let me be blunt: Nazis are bad. Period. They should not be making up a straw man argument just to appear balanced.
Finally, there's the entire clickbait news issue. Some news outlets intentionally post controversial topics in order to generate clicks and views for advertisers. SurfSafe has no method to represent issues related to bad reporting or ulterior motives.
The domain "getsurfsafe.com" was registered last month, on 2018-07-03. There are various domain reputation reporting systems out there. One of the big red flags is a new domain name. So this is a red flag. Nothing in their domain name or SSL certificate identifies the owner of this service. (And thanks to the GDPR, there's even less information available.)
How about their web page? They don't list any of their people and there's no "About Us" or "Who we are". The only thing it says at the bottom of the page is that it's powered by Robhat Labs. Robhat Labs is almost as mysterious: no names, no about us, nothing identifiable.
In fact, I found more information about the people behind Robhat Labs via LinkedIn and their press release. When it comes to trustworthiness, I have a lot of issues with site owners who don't identify themselves on their sites.
Sadly, some groups try to take advantage of this void and fill it with bogus analysis tools. SurfSafe is not a solution for identifying fake news.
The only thing worse than fake Fake-News detectors are news outlets that want us to trust them, but promote "new tools" without vetting them first. (I'm talking to you, Wired, BoingBoing, and Mac Observer.)
Thanks to R and M for pointing this app out to me. And thanks to The Boss for the lively discussions.
Lots of people want a one-button solution. Something that will quickly tell them whether a picture is real or fake. Sadly, there is almost never a simple yes/no answer. But this desire for a fast, simplified solution opens the door for lots of snake-oil solutions and charlatan products.
This time, the product is called SurfSafe. This is a brand new Chrome plugin that claims to catch fake news by spotting fake photos. News outlets like Wired, BoingBoing, and Mac Observer recently touted this wonderful, new plugin. (I can only assume that none of those reporters actually reviewed the product before summarizing the PRweb press release.)
TL;DR: In my professional option, stay away from the SurfSafe plugin. It has serious privacy violations and does not do what it claims.
How it works (in theory)
According to their web site, snazzy video, and the application (when you install it), SurfSafe helps identify pictures associated with fake news. First, you select the news outlets you trust from a long list of sources. Everything from ABC, NBC, and BBC to Breitbart and Fox are listed. SurfSafe appears to make no assumptions about the validity of information put out by the various sources.After that, you start surfing the web. Every picture gets a little icon above it. Clicking on the icon queries the SurfSafe servers for information about the picture. It returns the number of sightings, links to the sightings, and number of people who have reported the image. (Reports identify the number of people who think it is fake.)
Hovering the mouse over each picture also reveals a "Report" button. This is where you can report a picture as being "Propaganda", "Misleading", or "Photoshopped". This way, other people who see the same picture can quickly see how it has been reported.
Privacy issues
The first significant issue I saw is related to how SurfSafe gathers pictures for their collection. I thought that they would just harvest pictures from news outlets. I mean, they have a huge list of news outlets that you can select from as authoritative sources. However, that doesn't seem to be how it works. I went to a couple of news outlets on their list, and most of the pictures were not already indexed by their system.Then I looked at my web browser. With Chrome, you can press Control-Shift-i and pull up the developer panel. Then you can click on the "Network" tab and see all of the network requests. Here's what my FotoForensics service normally looks like:
Normally, there are 5 network requests: the web page, style sheet, banner picture, fonts, and my little IPv4/IPv6 experiment.
And here's what happens when I have SurfSafe enabled:
Every web page gets passed through the SurfSafe extension. However, that final 'query' is the troubling issue. It submitted the URL for the banner image to SurfSafe. If my web page has 5 pictures, then all five image URLs would be sent from my browser to SurfSafe.
The number of privacy issues here is stunning:
- The folks at SurfSafe know every web page I visit, when I visit it, and every picture on every web page.
- If any of the URLs contain personal information, such as access tokens or keys or names, then they get that, too. Since I tested against my FotoForensics web site, I watched the logs for their accesses. They definitely keep and use all URL parameters when retrieving pictures.
- If any of the URLs point to personal pictures -- things that you don't want other people to see -- then that's too bad. SurfSafe retrieves the pictures and adds them to their collection of known sources. Remember: lots of URLs are not publicly known but are still publicly accessible if you have the right URL-based parameters. For example, Facebook pictures can be marked as private. But the URLs are accessible if you have all of the URL parameters. This means that SurfSafe can access your private Facebook pictures as soon as you visit your private Facebook page.
- As you'll notice from my screenshots, I accessed my site using HTTPS. HTTPS is supposed to be secure, but it's not secure from local browser extensions. SurfSafe grabs your secure URL and passes it to their third-party service.
For my own testing, I visited a couple of web pages on the public FotoForensics site. Then I queried SurfSafe about my banner. Here's the result:
My banner logo was found on three sources. Those are the three web pages at FotoForensics that I visited over HTTPS. (I checked my logs: SurfSafe only visited those pages after I went there with my browser.) Granted, on their public "sources" list, they removed the URL parameters from the source URLs (meaning that any URL requiring parameters will be a broken source if you click on it). However, they still list the web page and the page's HTML title; the title alone may give away too much personal information.
In any case, if you use the SurfSafe plugin, then you are feeding the URL of every picture you come across into their service for their collection. This is a huge privacy issue.
Accuracy issues
To test SurfSafe's accuracy, I needed a sample picture. I searched Google for Supreme Court wannabe Brett Kavanaugh and his Wikipedia page came right up. On the Wikipedia page is his picture. According to SurfSafe, the picture had zero sources and zero reports.I uploaded the picture to FotoForensics. The metadata clearly states that the picture had been processed using Adobe Photoshop CS3. So, I used SurfSafe to 'report' the picture as 'photoshopped'. I reloaded the page and now there is one report for the picture (mine). However, nowhere does SurfSafe mention the cause of the report. There's just "Reports 1".
Clicking on Kavanaugh's Wikipedia picture takes you to Wikimedia Commons, where another copy of the picture is hosted. And when I say it's another copy, I mean it has the exact same SHA1 checksum -- bit-per-bit, it is the same picture. At Wikimedia Commons, SurfSafe reports 4 sources: 3 at Wikipedia (which appear to be the same URLs) and one at Wikimedia. However, the Wikimedia page doesn't have any "reports".
I then went back to the Wikipedia page and reloaded. Now it has 3 sources listed and one report. I waited about a minute and reloaded again: 4 sources and still one report. So even through a photoshopped picture is a photoshopped picture wherever it lives, SurfSafe only tags the photoshopped report to the one URL where I reported it as photoshopped.
What this means: SurfSafe won't flag a picture as being propaganda, misleading, or photoshopped unless someone else already flagged the picture at that specific URL. And since many web sites, including Facebook and Twitter, use user-specific URLs, it's very possible for other people to see the same picture and not be told that it's been altered. For SurfSafe to claim that they help detect fake photos is misleading at best.
Crowdsourcing issues
Having humans evaluate pictures and categorize them as real/correct (not reported) or propaganda/Crowdsourcing usually results in the most popular solutions. Consider the case of Boaty McBoatface. In March 2016, the Natural Environment Research Council ran a poll to name one of their ships. Through the miracle of crowdsourcing, the name "Boaty McBoatface" was selected. (This created a controversy when the NERC decided to assign the name to a different vehicle -- one with a less public image.) However, this shows how crowdsourcing works: it's not always the best solution.
There are other issues with this crowdsourcing solution. For example, nobody vetted me. I didn't have to login or create any kind of unique ID to submit my 'report' to SurfSafe. This means that anyone can poison pictures at SurfSafe. It doesn't take much for a bot to submit thousands of reports against lots of pictures. (And you can always use a list of open proxies if you want to come from ten thousand different addresses.) Organizations with specific agendas can easily have real pictures marked as untrusted in order to cause confusion or attempt to discredit negative publicity.
Perhaps SurfSafe should consider a Facebook-style user rating system. Facebook users who forward known-false stories will earn a lower trustworthiness score than people who forward vetted stories. This reputation management system is Facebook's attempt to combat fake news. The idea is that users with low trustworthiness scores won't be able to propagate new stories very far. If you have a reputation for promoting false stories, then any new content from you will be treated with a healthy amount of skepticism. Then again, SurfSafe will still need a way to vet stories before assigning user trustworthiness.
Trusted Sources
One of the first things that the SurfSafe browser extension requires is a list of trusted news outlets. Users select from a preset list of known news entities. At that point, it assumes that pictures from these sources are considered safe. This shows up in the JSON code seen in the network query. I marked CNN as a safe news outlet and then visited CNN.com. Each of the bulk query images uploaded to SurfSafe was automatically marked with the classification of "Safe".The problem is that no news outlet is perfect. Trusted news outlets may occasionally make mistakes. (Some news outlets have more mistakes than others. Right, Fox?) This is particularly important in this era of rapid news coverage; for many news outlets, speed is more important than accuracy. Without vetted news reporting, we end up with situations like the Boston Marathon bombing, where the wrong person was named by some media outlets.
There's also a false equivalency and false balance problems that some media outlets experience. A few journalists think "fair and balanced" means that they must provide equal coverage to both sides of a reported topic. If they say something negative about Nazis, then they must also find at least something positive to say. Let me be blunt: Nazis are bad. Period. They should not be making up a straw man argument just to appear balanced.
Finally, there's the entire clickbait news issue. Some news outlets intentionally post controversial topics in order to generate clicks and views for advertisers. SurfSafe has no method to represent issues related to bad reporting or ulterior motives.
Attribution issues
Okay, so now we know how SurfSafe works (or doesn't), we can start looking at who is behind this service.The domain "getsurfsafe.com" was registered last month, on 2018-07-03. There are various domain reputation reporting systems out there. One of the big red flags is a new domain name. So this is a red flag. Nothing in their domain name or SSL certificate identifies the owner of this service. (And thanks to the GDPR, there's even less information available.)
How about their web page? They don't list any of their people and there's no "About Us" or "Who we are". The only thing it says at the bottom of the page is that it's powered by Robhat Labs. Robhat Labs is almost as mysterious: no names, no about us, nothing identifiable.
In fact, I found more information about the people behind Robhat Labs via LinkedIn and their press release. When it comes to trustworthiness, I have a lot of issues with site owners who don't identify themselves on their sites.
Misleading at best
So let's look at what SurfSafe really does and compare it to their claims (on their web page):- Claim: SurfSafe protects users from misleading photoshopped and fake news throughout the Internet.
Fact: Nope. At best it finds other sources for the picture. It relies on unvetted people to flag pictures as real or not, and those flags do not carry over to other URLs that host the exact same picture. Even as a search-by-picture service, you're better off using Google Images or TinEye. - Claim: They mark the level of safety of an image or article on the corners of the pictures.
Fact: Nope. They add widgets to the corners, where you can then query their service. The "level of safety" is based on crowdsourcing and not any scientific analysis. Moreover, their crowdsourcing approach is significantly biased and easy for attackers to manipulate. - Claim: They defend the Internet and bring the world one step closer to a fake-news-free Internet.
Fact: Nope. They violate user privacy by collecting URLs and pictures that may be sensitive in nature. Not only does this service not identify fake news, it permits attackers to mark real news as fake. - Claim: You choose who to trust by deciding which organizations to base the truth off of.
Fact: You can select organizations that you think are trustworthy, but their selection process does not determine accuracy. Regardless of whether I think the BBC is good journalism or you think Fox is trustworthy, this says nothing about whether the reporting is actually accurate.
This is one of the big problems a silo mentality. Everyone is already using algorithms to show you what you will probably want to see. Facebook shows you similar articles to ones you like. Amazon shows you similar products that you might like. Twitter makes recommendations about who to follow based on your previous actions. This doesn't mean that SurfSafe is showing you the truth; it's only showing you what it thinks you want. - Claim: "SurfSafe uses image and textual analysis to catch when fake news tries to mislead you."
Fact: Nope. This may be what they think they are doing, but it is definitely not what they are doing.
Sadly, some groups try to take advantage of this void and fill it with bogus analysis tools. SurfSafe is not a solution for identifying fake news.
The only thing worse than fake Fake-News detectors are news outlets that want us to trust them, but promote "new tools" without vetting them first. (I'm talking to you, Wired, BoingBoing, and Mac Observer.)
Thanks to R and M for pointing this app out to me. And thanks to The Boss for the lively discussions.
Read more about Forensics, Image Analysis, Mass Media, Network, Politics, Privacy
| Comments (5)
| Direct Link

Cheers!
In the networking world, there's a technical difference between a bridge, router, gateway, hub, and switch. However, media coverage and product labels often use the words interchangeably. Your "switch" may actually be a router, and your "router" may actually be a gateway.
By the same means, plugins, extensions, and addons have very specific technical meanings. And you are correct: it is technically an extension.
However, news coverage called it both an extension and a plug-in. They don't use the technical definitions; they use the terms interchangeably. Since most users are not technical, they don't care what term since they are all just things that are added to the browser.
Since I'm criticizing the media coverage, I decided to stick with their terminology rather than the technical definitions.
> Using Chrome
Choose one. (Chrome/Chromium/Chrom* is a by default keylogger that sends everything you type in the address bar to Google search, how about making a blog post about that? It should be condemned and lit with Fire, and replaced with Firefox ;0 - ps: use Firefox nightly with WebRender enabled to do the compositing on the GPU it will give you +60FPS browsing experience with the right GPU, in addition you have the combo privacy.resistFingerprinting and the one about first party isolation which Chrome will never have)
I'm not ignoring it. I just don't know if it will be my next blog entry, or sometime later. I don't want to announce a specific blog entry topic until after I've written it.