Facebook Censorship
Following the exposure of Facebook's content filtering rules, Peter, Chris and Fraser discuss censorship and ethics in data science.
For more information on Aleph Insights visit our website https://alephinsights.com or to get in touch about our podcast email podcast@alephinsights.com
Transcript
Hello and welcome to the Cognitive Engineering Podcast produced by Tell Me Studios for Aleph Insights. In this series of podcasts we take a look at interesting topics and discuss what we think they tell us about analysis and decision making. I'm Fraser McGruer and I'm here with Peter Coghill and Chris Wragg of Aleph Insights and this week we're discussing Facebook's censorship of certain content. Okay, great. So Chris, can you lead us in please?
Speaker B:Yeah, so this really boils down to a series of revelations published in the Guardian newspaper to do with Facebook's set of criteria for censoring material which is published through Facebook and it's given such sort of revelations as, you know, how Facebook makes decisions about things to do with, you know, when is it acceptable to show images of children being bullied for example, you know, whether it's in a context of somebody sort of propagating that bullying or whether it's actually somebody drawing attention to it for sort of information and bullying prevention purposes. So trying to distinguish between those same content but different context and a whole series of other sort of kind of specific social media ethical issues which have arisen. Things like how to deal with revenge porn and some of these other modern phenomena. So really our interest in this is first of all the technical challenges of determining things like context and how you decide whether content that is the same is okay in one context versus being unacceptable in another context. But also just more broadly, you know, some of the issues it raises about ethics within data science and, you know, I think this is a growing area which we need to think about. So how does Facebook do that? So
Speaker A:how do they police this and how do they decide, as you said, if it's similar content, let's say taking the example of bullying, how do they identify whether it's just people being horrible, showing some bullying or showing, you know, raising awareness? How do they
Speaker B:do that? So I mean perhaps Peter can talk a bit, you know, technically about how it's achieved in a moment but in terms of the sort of rules that they follow or the general principles that they apply to deciding things like this, this is what's actually being published. So for example in the context of, you know, images showing a child being sort of physically abused, if there is kind of celebratory language around that or it's, you know, gloating in its context then that's, you know, that's the sort of criteria they would apply to removing that content. Whereas if it's, you know, in the context of isn't this a bad thing, you know, here's an anti-bullying advert, for example, then that's seen as permissible. So it's about the context of the language that's being used around it and also the other kinds of content that surround it in terms of, you know, if it's a whole series of videos of children being bullied or images of children being bullied then that might ring some alarm bells.
Speaker A:Okay. Peter, so how do they do it? How do they?
Speaker C:Well what's quite interesting is that what The Guardian has exposed is there isn't a sort of supercomputer that understands everything that people are posting on Facebook, able to sort of scrutinize images and work out what they mean and look at words and understand what they mean. It's largely driven by human workers and I think Facebook has around 7,000, although they have plans to have more, who sit down and read posts and they, I think they're probably anonymized in some ways so they don't know who the post is from, but they read posts, they look at videos and they look at images and they apply their own judgment based on a set of principles laid out in terms of what is appropriate and what is not appropriate for various pieces of content. So in a way it seems, you know, it seems kind of a bit old-fashioned that they have these, you have these arbiters of appropriateness, human agents applying simple rules and it's probably, I imagine it's assisted by machine learning techniques so once a post has been deemed inappropriate then the system will recognize the same post because an image match
Speaker B:is quite an easy thing to achieve. There's also the reporting mechanism of course that, you know,
Speaker C:Facebook. Yeah, so and I think, yeah, and the two billion users on Facebook can self-report when they find something inappropriate or offensive and that presumably queues up things to look at for the 7,000 Facebook workers to look at. So yeah, so it's sort of interesting because it's so technical, it's not very technical and I think this tells you something about how difficult it is to do this sort of thing. You still need to rely on humans to interpret it and then make a value call on the content and we haven't yet got systems that can automate that. Now I think that raises the question, well do we want to hand over to automated systems to decide what is appropriate? You know, do we want to sort of take away censorship? You know, do we want to hand over censorship to machines and let them roll with it? Because that will be a world where rather than humans checking up on other humans, you'll have machines checking up on humans which means that we'll hand over control of what is ethically correct to another type of intelligent agent other than humans.
Speaker A:Well this reminds me of when we had, and his name escapes me, when we had the chap in talking about fraud and was it cyber fraud? I can't remember but it was talking about fraud. Cyber fraud. And one of the things he was talking about is actually it's quite a largely automated system where there are certain behaviours that can quite easily trigger an alarm. But ultimately, if I recall correctly, he said the ultimate judgement call came down to a handful of analysts who would sort of make a yay or nay call as to whether this was fraud or not. Is there, I mean, I'm thinking about the language that's used in posts of this kind for example. Is there, and you talked about how it's relatively simple for image recognition etc. What's to stop Facebook from adopting a similar approach? Well I'm sure they do but ultimately
Speaker C:it comes down to human again to make a call. You can't, at the moment it seems that even with the might of Facebook's engineering prowess they haven't been able to eliminate
Speaker B:the need for a human arbiter. And I think the case that really flags that up is the Vietnam War photograph issue that came up where it shows sort of Vietnamese children who've suffering during the war and one of them's naked. And the photograph was initially banned because of their policies on child nudity. And something like that is a very complex call to make. And I think for me the interesting thing is that actually it's not a question of should we hand these things over to machines. I mean I think you know Peter, the maths that Peter gave earlier on of you know two billion Facebook users and 7,000 people looking at things. If you simply consider the amount of content that 7,000 people can actually review and the amount that's generated by two billion people then that's just not a sufficient resource to deal with that. And you might talk about okay using initially the machines to queue the information to the human analysts so that they're not looking at everything. But even so you know the number of calls they're going to have to make is going to you know cognitively overload them I suspect. So I suppose the issue is not you know should but how are we going to do this. And in any case when you have humans doing that they're essentially following written down algorithms when they go through this set of criteria that you know a decision tree that's been written for them by Facebook and then they make an opaque interpretation of that anyway. So it's not like we can open up their brain and see how it how it is they've made the judgments they've made to ban or not to ban and there's no particular transparent audit process for that. So I don't see as divesting that to machine learning is really a great leap to be honest.
Speaker C:Because you're still trying to systemize it no matter what you do even if it's humans pressing a button. Yeah I mean it's sort of a similar model to crowdsourced work where you're slotting into a process of turning information from one form into another or making a judgment about information and rather than a sort of machine module you've got a human module in there and harnessing the highly parallelized and extremely clever algorithms that work in a human mind.
Speaker A:Okay so again is this a case of well case closed because if we look at a couple of the issues or concerns people might have one of them it's kind of a how because this is a question of volume and so the question and the answer is volume. Why is this such an issue or how do they manage it or just because it's so big and I guess going back to my example of fraud I guess maybe they do run it in exactly the same way as a fraud detection company might
Speaker C:it's just the numbers are different right. And the question is somewhat different I think because fraud is a whether or not something is fraud is a very narrow part of a similar ethical framework but the rules are much more concrete. And the data around it is much more quantified. So yeah it's a similar but easier problem I think. I mean this is essentially a really
Speaker B:difficult sentiment analysis problem I think in essence because you've got you know what is what is the sentiment around this particular content so you've all you know perhaps the most difficult problem is not identifying the material which may be you know maybe it should be subject to banning or not but actually around making the decision about whether or not it
Speaker C:should be banned. And yeah there's a bit of decoding what the intent of the poster was because they might be posting a sort of image of something as a as a counter example of something that shouldn't be allowed so here's a terrible image let's stop this let's ban together if there's a
Speaker A:force for good to prevent this sort of thing happening. But even with that I mean I think there is a way to ultimately quantify that stuff because you can have a you know a sentiment score or you would have to establish some kind of rate card right. So but it's a very it's a very
Speaker B:difficult problem because you know humans use irony, sarcasm, hyperbole you know double entendre and we misinterpret one another all the time so you know and we're very adept at using natural language whereas you know a machine learning algorithm to try and encode and then you know teach it how to interpret that kind of thing is very difficult I think. I think
Speaker C:what something I think one interesting aspect of this is the media's comment on what appear to be overly simplistic rules so one of the things that the media is presenting it in the case of oh isn't this a bit shocking because these rules seem so simplistic when actually I think that Facebook has developed these rules over many years since it's been around and it's in its own interest to make the best rules it can and the rules by necessity have to be quite simplistic because you have to pay you have to employ people to work fast and to make and empower them to make sound judgments very quickly so I think that you know it's probably about as best way of doing it as you can with human agents and that's not to say that it could be improved it couldn't be improved with technology and I'm sure I've no doubt that Facebook is doing everything they can to improve the improve the system because it is totally within their own interest to provide
Speaker A:a better better service. I mean presumably one would want to err on the side of caution here talking about that it's slightly complex and you'd want to err on the side of caution because I guess going back to the fraud example you know it's more clear-cut has something illegal happened yes or no and the downside is if you get it wrong is the I know someone gets away with a fraudulent act for example whereas this I mean let's say you erred on the side of caution and stopped more posts than you might normally the downside is just get angry people yeah well it's a classic sort of
Speaker B:type one type two errors problem of you know do would you would you would you rather let things go through or would you rather stop things and you're thinking about the vault you know and also you know what is what is Facebook's censorship role so you know I mean the platform is used and they're in a competitive market and if their if their platform suddenly stops oh I can't I can't post my content onto Facebook they become overly authoritarian or too sensitive then people will migrate to other other platforms and you know they were set up as an open communication you know with an open communication ethos and not to get engaged in unnecessary censorship so I think it's a it's a setting that sensitive you wouldn't have to you wouldn't have to have a particularly low threshold to make a huge difference to the to the number of posts given the volume of of you know content that gets gets posted to Facebook so yeah I think Chris touching
Speaker C:on an important point there it's like well they this is censorship right and censorship historically has been used for both good and bad things but censorship in you imposed incorrectly or inappropriately can drastically alter the way that the the public mind thinks and what is important in the public mind so Facebook is having to tread a very careful line and not notwithstanding the fact that they are this is a global community this is with with all the diverse cultural backgrounds and cultural norms and what is it what what is appropriate ethically appropriate for different cultures there is no world culture where everyone agrees and we're all sort of moderate liberals and you know we don't we all agree what is appropriate and what is not they're having to balance it for all different types of all different parts of the world but I think wider though why do we expect Facebook to have this role why do we look to Facebook to self police on this we don't expect paper manufacturers to keep tabs on what everyone is printing you know people can print child pornography on on paper and we don't go to Banner and and complain that they've been allowed to buy that child pornographers have been allowed to
Speaker A:buy Banner paper okay but I think there are a slightly provocative question because I think even I suspect you probably don't quite believe the example that you've just said but we'll answer the question I mean do you expect Facebook I honestly don't think that Facebook should have
Speaker C:any role in self in being in charge of censorship I think it's almost dangerous to allow the platform owner the platform creator to be in charge of deciding what people say and think if Facebook their mission is to provide a an open platform for people to connect and communicate that's their main mission so really it's up to the users it's up to the society of people who use it
Speaker B:to use it appropriately is that realistic well I think I mean I take a slightly different view in the you know my my view is Facebook are effectively a media company and we do expect you know broadcasters for example to be responsible for the content even if it's a news interview if the person they're interviewing starts swearing we expect the the news in we expect the broadcaster to be able to cut from that broadcast or to you know censor it in in some way and you know in other Facebook does make make money of course from from offering this this service you know that it is a commercial entity and therefore I think they have an obligation you know much in the way as a train company has an obligation to prevent its passengers from coming to harm by the train crashing or allowing you know some somebody to be abusive on the on the train I think we can you know we can expect Facebook to make reasonable adjustments for that to to be the to be the case and also they are going to be best placed to do that they understand the technology best and they are sort of at they are at the cutting or they are at the coalface of where this material is being placed and um you know have uh access to the the way the way it is posted in the in the first place so they're they're they're basically best placed to be able to censor whether they should be holding themselves to account for that censorship is another matter whether you know there should be an objective external body that sets the rules and then you know Facebook are charged with uh sort of policing that or or or implementing it I suppose is is you know perhaps another model but I do I do feel they have an obligation I think I I can only imagine the company itself feels they they also have yeah I think they yeah I mean like I'm in my vision
Speaker C:they have they offer an obligation to allow society to self-police um so that that you rather than have rather than having as employees of Facebook and therefore being biased by Facebook's um desire to make money so you know one of the one of the reasons one of the drivers for Facebook policing content is so that they don't get their sponsors i.e people advertising on Facebook next to unhealthy content because that's bad for revenue it's not in it that that particular particular dimension is not an ethical thing it's not an ethical driver it's a that's purely profit driven where a better system perhaps would be to allow independent bodies of people to police Facebook and Facebook just provide them the tools to do that much in the same way that sites like Wikipedia do they have a set of principles and a sort of collective mission and it self-polices um it's not there's no single arbitrary body that that uh yeah I mean I think
Speaker B:obviously you know community organizing things like you said you know um ethics are are relative and you know what what happens if you get a particular you know if you if you had a sort of Mary Whitehouse kind of um movement you know develop on on something like Facebook and very heavy censorship uh you know that that would be a would be an issue but like you say yes that that that commercial and ethical conflict of interest is it is an interesting one okay we need to wind
Speaker A:up before we do is I feel that we're just starting to get into something here it's yeah um before I wind up is there any final points that anyone wants to make looking at Peter yeah yeah it um
Speaker C:this all touches on um the broader issue which we we were trying to use Facebook as an example to get into this topic but just in various in summary but data analysis as a as a thing is an enormously powerful and amazingly useful tool that we've generated as humans as a as a as a technology and it has great force for both good and potentially misuse for bad for evil um but there's a certain part there's certain sort of features of data analysis which can obscure the truth and can have to be looked out for as if you're a data scientist and just just just a few I mean the data is sort of uh it often held up as providing a having an inflated level of objectivity just because it's written in hard numbers and you've got an algorithm or a model which helps you predict something doesn't mean that it's it's truth right there is no truth out there there's only a there's only uncertainty and information helps you erode that uncertainty so it's not just truth because it's numbers um it there's also with with data often unintended proxies creep in and what I mean by that is if you're if you've got a large data set about about people and you're trying to make uh predictions or policy decisions to to help those people in some way um but your your your your ethical framework precludes you making any judgments about based on race or maybe background or age or some sort of rule uh so you eliminate those parts of the data so you say well we're not gonna look at age and we're not gonna look at ethnicity um there may well still be proxies in the data uh that that act as proxies for race so even the although you're trying to exclude it where they live might be a very strong predictor of what what their ethnicity is so uh it's very difficult to exclude those things so that's something else to watch out for and another key thing is the phenomenon of compounded assumptions so at every stage when you're conducting analysis with data you're introducing assumptions maybe maybe explicitly and maybe well recorded but also implicitly in what your selection of what tool you do into the section of what method you use uh and or the way you clean up the data the way you process and the way you visualize the data all add assumptions and which cloud further the potential truth as such as it is within the data so as a data scientist you have to be cognizant of the limitations of your of all methods you're using and and i think
Speaker B:we we are at a we are at a juncture where data science has um uh you know is a is a sort of blooming um sort of area of science and technology at the moment and uh but we are creeping into the territory where um you know some of some of these companies that have allegedly been involved in um uh data science um exploitation of um elections and trying to uh do micro targeting of of swing swing voters um and uh you know whether or not there is truth to that we are in danger of um a situation creeping in where data science that we allow a narrative to be created where data science is seen as in some way uh nefarious or um uh you know uh opaque and mysterious and conniving and um i think you know data science data science the data science community need to get ahead of that curve and demonstrate that they are thinking uh about the ethics of data science and that they are doing things with transparency and humility um and i i think this is a an area that needs to be um you know much more robustly debated it's a shame because i feel that almost
Speaker A:now our podcast should just be started because i think we're just starting to get into some stuff here so we'll have to park it and save it for another time very interesting um before we sign off um you're on facebook um we're friends on facebook right i'm looking at peter um chris are you on facebook uh i'm not i thought you might not be so um just very very totally off the grid
Speaker C:yeah he may as well not exist there'll be no record of chris after his death ever
Speaker B:yeah anyway just me in my uh in my apocalypse bunker yes quite so just very very briefly
Speaker A:because we don't have a huge amount of time what do you mainly post on on on facebook peter very
Speaker C:little uh but it's usually some it's usually it's usually something i found funny uh i'm not a big fan of posting endless photos of me at posh restaurants or on holiday and things try to i try you have plenty of opportunity to do that which i guess of course but uh i i desperately try not to fall into that trap of projecting a false image of me uh on facebook i like to try to use it as a uh provide my friends and contacts with a sort of bit of a window into my soul by exposing those things which i find particularly funny um or amusing okay that's a frightening
Speaker A:window yeah i was about to say that's like yeah i'm yeah i yeah i i shall look because i don't see your your post no i don't post very often yeah because i use mine rather boringly like a lot of people is just to show the world how wonderful my kids are and attractive and you know it's pretty boorish it is it is rather i admit it yeah um so chris you must be nice this
Speaker B:101 times so why aren't you on facebook uh well that's that's a podcast in itself okay let's let's
Speaker A:not go down that that rabbit there's a very big secret there yeah yeah that's right okay gentlemen i really really enjoyed that thank you very much i'm frase mcgrew we've been here with peter coghill chris wragg um listening to the cognitive engineering podcast with aleph insights thank you and until next time goodbye
