Terrorism Prediction with Big Data
Fraser, Nick and Peter consider the use of Big Data for prediction of terrorist attacks.
New online ecology of adversarial aggregates: ISIS and beyond: http://science.sciencemag.org/content/352/6292/1459
For more information on Aleph Insights visit our website https://alephinsights.com or to get in touch about our podcast email podcast@alephinsights.com
Transcript
Hello, and welcome to the Cognitive Engineering Podcast produced by Tell Me Studios for Aleph Insights. In this series of podcasts, we take a look at interesting topics and discuss what we think they tell us about analysis and decision making. I'm Fraser McGruer, and I'm here with Nick Hare and Peter Coghill of Aleph Insights. And this week, we're examining a new model for forecasting ISIS attacks that uses social media data. So, Nick, sketch this out for us. Tell us
Speaker B:about this new model, please. Yeah, this is about a paper that was published in Science recently, written by Dr. Neil Johnson at the University of Miami and his team. He's a physicist and has developed an approach to analysing ISIS social media activity that actually doesn't quite do what the headlines suggest it does, which is that the way it's been reported is that it will help us predict ISIS attacks. And if you look at the paper, which is not easy, it's not a straightforward thing to follow. It's not actually totally clear exactly what it was that they were doing and how they were using the information that they had. But as far as I understand what it's doing, it's trying to look at the dynamics of ISIS activity online, and with a suggestion that those online dynamics correspond to real world events. So, the approach they've used is something very interesting that certainly would not have been possible 10 years ago. So, effectively, what they do is they look for these online, what they call aggregates. And an aggregate is something where you have on a social network, so whether it's Facebook, or in fact, they mainly use data from Vcontact, which is a Russian social media platform, which is not as routinely interdicted by law enforcement authorities as Facebook is. And they look for, basically, clusters of followers who all followed each other. So, eventually, they would start with one clear node, which was a pro ISIS node, look for followers of that node, look for followers of those followers, and identify these emergent clusters, these aggregates. And what I think they demonstrate, but not very clearly, but I think the suggestion is that when these aggregates have a dynamic of their own, so they start small, they cluster together, they build up. And the suggestion is that when there's a sufficiently large cluster, when the aggregate is sort of big enough, that that is a predictor of an attack. And they generate some predictions from that, they suggest that actually you might be able to disrupt terrorist attacks by stopping these clusters growing by interdicting and shutting down those clusters before they have a chance to get big enough. It's, you know, they say that they've identified, using this method, 196 pro-ISIS aggregates, with a total of about 100,000 individual followers. And, you know, that tells you straight away that what it isn't going to be doing is predicting ISIS terrorist attacks, because there have only been a handful, at least, in the West. And, you know, there haven't been something like 200. So, but, you know, it's an interesting approach anyway. And I think it, you know, it's interesting in that it tries to do something which is, has only recently become possible. So, sorry, is it meant to predict anything at all? It's meant, there's the suggestion, and again, I don't, this is not, I don't think the paper demonstrates this, although they may well have done this. But the suggestion is that the emergence of these large aggregates is a forward indicator of an attack. But what it definitely doesn't do, it's not going to predict a total lone wolf, somebody who has no presence on social media. And it also isn't going to tell you where these things are going to happen. So it's not, you know, it's a limited use from a law enforcement, from a tactical point of view. But from a strategic point of view, you know, if we were thinking, well, when do we need to start worrying? When do we need to step up our security? You know, it has the promise of providing at least some more evidence for that. But I think the reason we want to talk about this today is more the general theme of, you know, how do we, of using social media data, and which the data set is very big, to draw out conclusions that in the past we would have relied on, you know, policy experts to do.
Speaker A:Great. Peter, pick this up and run with it.
Speaker C:So, well, I think it's interesting, it's useful to consider what social media can tell you. So if you're designing a data collection scheme as part of a system, you usually have a specific question in mind. So you tailor the data collection to suit the question you want to answer. And this is a reverse situation where the thing exists, and the data is there for a particular purpose. What else can it tell us other than what it's there for? So social media data generally is about the social networks that it supports. So it's who knows who, who converses with who, who likes what things. So, and it's often harvested for deciphering people's preferences for marketing campaigns. So your Facebook stream is sensitive to the things that you have liked in the past, potentially the things that you have clicked on and dwelt on, and it will show you more of those by preference than things that you've skipped over. So that's kind of what it's being used for. So that potentially has useful utility in defence and security. So if you have people who like bad things or things you consider bad, then you can use the same data to weed out who those people are, and then to sort of see what the associations are with people and their preferences. And likewise, you can see who's talking to who. So I think that's the generally the methodology that this paper refers to. It's these bubbles of concern, give you an indication of where the where things are like, where the problems are likely to come from. So but I think what, as Nick said, it's not going to help you predict the lone wolf. And it won't help predict any organisations with mature operational security. So people who don't talk online or social media who are who you know, the spy agents around the world are quite mature in this and they have their encrypted, encrypted communication channels. So you won't pick up much stuff in the data. However, it there may be holes, there may be gaps in the data. So if somebody is, is not using social media at all, that might indicate to you that they're there's something odd with them. If, if they're the sort of person you'd expect they were, but are choosing not to for some reason, then there's there's potentially there's the lack of data can tell you
Speaker A: ion. I think this was back in: Speaker B:not, I don't know. There's a very famous actually example, it's often used as a cautionary tale, I think probably very unfairly, but of Google flu, which did a very similar thing to try and predict where there were going to be flu outbreaks. And it essentially looked at flu related search terms. So people who were looking up, you know, symptoms like having a runny nose or a fever. And when Google ran this on the past data, it seemed to produce quite a robust predictive model. But then the there was a sort of structural change, which stopped it being predictive, which is that, you know, because Google actually started suggesting flu things, it got Google was good enough to give you appropriate search terms, the search data itself, then stopped, you know, this started to be a circularity in it, which which meant that the model stopped having the same degree of predictive usefulness. So that's the observer affecting the model
Speaker A:by ceasing to be an observer. Yeah, I mean, it's sometimes rolled out to say, oh, you know,
Speaker B:people who perhaps don't think or haven't don't really understand what big data means to say, oh, look at that big data can't can't really solve our problems. It's very unfair, because it's actually just it's quite a quite a specialized kind of case, actually, that sort of thing. I mean, by and large, you know, we what we have here is the issue where, and I think this paper is a good example, where you we now the information has long since stopped being in a lot of these cases, stopped being the constraint. The constraint is actually our ability to generate insights about what is creating that data. And here is an example where, you know, in this paper, they've developed a sophisticated model, a model that a machine would not be able to develop. So this is this is a model that is, you know, has been designed to try and capture what we think will describe social media activity. So sorry, the data is being used to test that. Sorry, why? Sorry,
Speaker A:can you clarify that why a machine wouldn't be able to generate this? Which is the general the
Speaker B:problem of sort of generalized model development. So if you have a reasonably complex system, where, you know, there isn't a very clear linear relationship between some values, some parameters, and output data, that the problem of sort of fit, generally just fitting a model to data is a really, really hard one that is nowhere near solved. You know, at the moment, there isn't anything where you can just you can give a load of data to a machine, and it will tell you what the most likely data generating process was that produced that data. So it's not that is something that at the moment, still requires human input, to be able to generate theories about what might be driving that data, what might be producing that data. And then the machine, what the machines are very good at is testing whether that theory is consistent with the data we have. Got it, Peter.
Speaker C:But I think it's, it's machine learning will generate statistical model, which is a model for the, the system is producing that data won't explain the physical processes within the system. But it's, it's, it's an approximation that you can then use to predict based on different inputs. And I think that I was, I was at a Google conference slash training session yesterday. And machine learning, as a thing you can, you can use is now in the hands of the consumer. There are there are other cloud platform providers available. But Google have an advanced, a mature system that for a small subscription paper use contract, you can you can design your own, your own systems to make use of their machine learning infrastructure to generate your own systems. So I think that, to me, that's an indicator that machine learning has reached a certain level of maturity and is mass democratizing these advanced computer science capabilities. And I think they'll only ever get better. Great. Let's have a little bit of
Speaker A:free consultancy then. So what does this mean for me? I mean, does this apply for me? And I know that you mainly consult to, I think, to large bodies, large institutions, to government and to commercial organizations. Okay. I don't think you consult to, to companies of my size, right? Maybe you do. I don't know. I, tiny companies. But so how could I use that? If so, let's say an issue I have with my film and photography business is, is, is building, is growth, essentially finding customers, essentially. So how can I go to Google and use their information, this sort of information you're talking about to get new customers and
Speaker C:to build my business? Well, I'll address that in a minute, but I can guarantee you've already used their capabilities. Oh, I have. Yeah. So if you've ever used Google Images and asked it to look up a picture of a dog. Yeah. I do that every day. That's using, that's you, if you look, or cats or kittens. No, no, I don't do cats. That, that, that, that is using the same infrastructure that you can now, you can now rent to, to do your own thing. So and if you've ever used the Translate, Google Translate or an amazing system, actually amazing system and more and more and more third party developers are going to start using this to provide new and exciting services. So you'll see that you'll, you'll, there'll be more apps available for built on this infrastructure to do, to do cool things. So what was your question of how you might grow your, your exposure or your market share? Good question. Let me think about that. Okay. Well, I can, I can,
Speaker B:I can at least something. Go on, make me a millionaire. Well, one of the big issues, one of the great unsolved problems in entertainment is what, why people, why, why some films are successful and some aren't. Okay. And why some books are, you know, what you, in general with these sorts of things especially in this day and age when entertainment could be consumed by anyone is you tend to have a few massive hits and a huge number, long tail of, of things that lose money. So what will be hugely valuable for you as well as for, you know, Disney and, and the other big studios is coming up with some way of predicting how successful a particular film will be. Um, that is the kind of thing, uh, there is a quantification challenge, uh, which is, you know, what actually, how do we quantify variables like how stirring the music is or how handsome the lead is or how convincing the, the, uh, story is. Those things are hard to quantify, but we, but we, you know, this opens the possibility that we can start to build models of those things. So, you know, if we want to predict how, uh, good film music is, we now, you know, have the computational power to be able to, to codify music. We, we, you know, to, to classify its shape and its structure and, and then to correlate that with, um, you know, things like listens or sales. Uh, so, so, you know, that it opens these possibilities up. Now we don't, we know in those real world problems, we're a long way from solving. Um, but, but, you know, we, we are, it's not impossible to imagine that happening quite soon.
Speaker C:So to work.
Speaker A:Hold on, I'll come back. Sorry. I know you're going to answer my question, but, but I suspect the studios and, and big music producers, record labels, et cetera, already do that quite a lot. You suspect wrongly. No, no, no, no, no. I bet they already do that. I think they do already use that data and they just, but this is a refining of it, I think, and that they might not make use of yet. But I think that's why you get so many average sort of, you get this kind of homogenization and these sort of summer blockbuster films because that's kind of what they're doing.
Speaker B:Yet another superhero. Yeah, exactly. Do we need another Spider-Man remake? Right. You know, the last ones have sold well, so let's just keep flogging that. Right, exactly. I mean, in a sophisticated model would take account of super, of Spider-Man fatigue.
Speaker A:There's a, yeah. God, yes. Spider-Man fatigue. I get that most nights. Um, there's, um, there's, you know, the, the, the book, uh, Data is Beautiful or something like that. Information is Beautiful. Information is Beautiful. Sorry. Have you seen that? David McCandles. It's a beautiful book and there's some really interesting bits in it. But one of them is they have three years worth of, of films and they plot it on an XY basis with budget versus box office, um, no, no, sorry, um, box office success, um, on the X, um, um, what's the word I'm looking for? Y-axis. On the, on the, Vertical axis. Yeah, on the, on the vertical axis and, and the Y one being, um, X one. Horizontal axis. The horizontal line, um, being, um, critical success. And then, actually, no, no, no, no. Oh yeah. And then, and then the size of the thing is budget. Yeah. Because there's a small, there's a bubble chart. Yeah. And then, and then the other thing is they actually colour them to what kind of story it is.
Speaker B:Yeah. To genre. Exactly. Interesting. Is that, I mean, cause, cause I, what is the correlation between, so let's, let's assume we can use something like Rotten Tomatoes, uh, to give us a sense of the, the critical, uh, claim and obviously box office is quite measurable. So what, what is the correlation?
Speaker A:Um, I think there's actually, oh God, I can't remember off the, I, I'd like to say,
Speaker B:Is it, is it clear or is it,
Speaker A:No, it's surprisingly unclear. Right. I mean, and what's really nice is every now and again, you see, um, you know, quite a small little, um, dots, whatever it is way up in the right hand corner, uh, which is nice. Yeah. Yeah. But sometimes it's quite satisfying to see a huge one that's way sort of off to the left.
Speaker B:Yeah. But what this suggests actually, which is actually quite tragic, I think is that as a, as a, um, if you are aiming to produce box office success as a studio, you, you probably can ignore trying to get critical success.
Speaker A:Peter, you were going to answer my question about how I can use Google, these Google tools.
Speaker C:So assuming that we can build you, uh, uh, a package of analysis using the, these machine learning platforms cheap enough to make it worth your while, what we could do is we could, we could take your, um, your current offerings, all the films you made and all the shows you produced crunch those and, and, and try to model what the common parameters are in that. So it might be, you know, if we're talking about film, it might be sort of the, the dark, dark and moodiness, the romantic core quantity, quantity, um, sort of the, the, the, the classifications of, of these things, of these, of these, um, productions, we could then, we could then track, we could then take the, the, the data that you've, you've, you've, you've tracked of who's watched your movies online, uh, and who listens to your podcast that you produce, uh, and, and see if there's a correlation between types of people and types of product. How long they watch for and when they switch off. Yeah, exactly. So we can, we can, we could, we could, we could, we could model what your, how successful your, your productions are. Uh, and that, that, that tells you a lot, that tells you sort of what type of people you appeal to. So you could market directly to those people that you, with stuff you've already produced, and it shows you what parts of the market you are not accessing. So it might, it would suggest you to tailor your artistic content to open up new markets. Yeah. If I could get you to do that for me,
Speaker A:I would do that. Um, very, very, very briefly, anything you want to say at the end there
Speaker C:before we finish? Okay. Yes. So Peter. So I think the, the, the, the, the analysis of people's sentiment and who they, uh, uh, associated with and what they say online, what they, you know, what they say they might do online. Um, I don't, uh, the, the, you have to be careful because that, that, that will suffer from similar things that polling data suffers from the disconnect between people's stated intent and then their actual behavior. Uh, there's a, there's a difference when you say, so we'll do, will you vote on next week to leave the EU? Um, a lot more people might say no than actually then go to say yes in the poll. How does that connect to the ISIS thing? How does that, I mean, if you're, if you're, if you run, if you're, um, looking at the associate, if you say just, just looking at these, these collections, these aggregates of people, um, and they are centered around, uh, um, particular clerics who are very, very anti-West or very, very pro-Islamic state, um, just, just by association alone, it doesn't necessarily mean they fully believe or fully commit to, to any particular, uh, direction. So that, that, that it may be bravado, young men just wanting to go along with the crowd and be the, you know, um, talk the big and, um, that, but it might, but they, some of them might be genuinely, uh, uh, quite
Speaker A:sort of, um, worrisome. Got it. Okay, great. Good note to wrap up on. Um, thank you chaps. Um, so explored all sorts of things there, but essentially around the theme of big data. Um, so I'm Fraser McGruer. Um, we've been here with the Cognitive Engineering Podcast with Aleph Insights with Nick Hare and Peter Coghill. Thank you very much for listening and until next time, thank you and goodbye.
