Episode 125

full
Published on:

26th Oct 2018

Future History

Why did it take 100 years to find the #49? In a world where everything is digitized, will we ever lose anything? If things don't have a digital record, will they may as well not exist?

Inverted Jenny: https://en.wikipedia.org/wiki/Inverted_Jenny Data created per day: https://blog.microfocus.com/how-much-data-is-created-on-the-internet-each-day/ How does the Wayback Machine work?: https://www.forbes.com/sites/kalevleetaru/2015/11/16/how-much-of-the-internet-does-the-wayback-machine-really-archive/#48383d419446 How much data is on the internet?: https://www.sciencefocus.com/future-technology/how-much-data-is-on-the-internet/ All the books in the world: http://www.gizmodo.co.uk/2016/12/every-book-ever-published-would-fit-on-to-one-hard-disk/

For more information on Aleph Insights visit our website https://alephinsights.com or to get in touch about our podcast email podcast@alephinsights.com

Transcript
Speaker A:

Hello and welcome to the Cognitive Engineering Podcast produced by me, Fraser McGruer, for Aleph Insights. In this series of podcasts we take a look at interesting topics and discuss what we think they tell us about analysis and decision making. I'm here with Chris Wragg, Peter Coghill and Nick Hare of Aleph Insights and this week we're discussing the discovery of one of the lost inverted jennies. Nick, first of all, what's an inverted jenny?

Speaker B:

's seen the gynecologist, the:

Speaker C:

By the guy who bought the original pane I think. Yeah, he bought it as an investment essentially. He ran away and hid from the post office workers he wanted them back.

Speaker B:

They've basically been over time, that pane has been broken up and sold in part to people but some of it, a lot of it's been completely accounted for and they know who owns it and some of them are on display and so on. But there were one or two which had been lost to science and no one knew where they were. They didn't know if they'd been destroyed or like in Brewster's Billions where he actually uses it as a stamp. Oh sorry, spoiler alert. Anyway, so this until about a month ago when what they think is inverted jenny number 49, which is one of the few remaining ones that hadn't been accounted for, seems to have turned up in someone's family vault. I mean, so great. The thing that interests me though is that the surprise is not that it's been found really that it's taken so long to find it given how easy it is now to find information and to find out about for example stamps and what they're worth. It seems to me that more and more things are appearing in digital form and I was thinking, well in the future is it fair to say that there will be almost no mysteries left or will everything be there on the internet to be found?

Speaker A:

Okay, you mean you're just starting to get into that, but can you elaborate when you say given how easy it is to find things now, just be a bit more explicit about what you're talking about there. I mean, you've kind of got into it. Well, I mean, yeah, I mean, so for example,

Speaker B:

once these inverted jennies are found now, their records are going to be on the internet, centralised and anyone who wants to find an inverted jenny in the future will look up on a list of inverted jennies and find out where they are and we will know for sure that, you know, we have them all. You know, that wasn't the case 50 years ago. There wasn't a centralised, easily found real-time database of these things and so it just seems to me, you know, as time goes on, more and more things will have a kind of digital receipt if you like and, you know, the question is in future, will all things be known? You know, if it isn't on the internet, will we just say,

Speaker A:

well, it's lost to humanity? So I think what we're talking about here is in the future, history should be simple, should be easy, right? Or is it? But also, the first question that sort of comes to me here is something we've discussed before. It's information versus, so storage versus processing, I think is an issue here. Who wants to jump in? Well, yeah, I'll go for it. So yeah,

Speaker C:

so yes, I think you're right there, Fraser. So storage costs, so the cost of storing information on a disk or a hard disk or a flash drive or whatever is pretty much approaching zero or will pretty much approach zero. The trick though is raw storage is one thing, which is very different from a curated, managed source of information. That's still going to be quite costly and expensive because that still involves people at the moment. You still have to have somebody who knows what stuff is where, or at least manages a catalogue of stuff that is well maintained so that future people can find things again. But technology, particularly sort of artificial intelligence, potentially could relinquish that because you could query all of the stuff and find what's salient to any given question. So yeah, it's likely that storage costs are going to approach zero and useful storage, i.e. stuff that you can retrieve things from, will approach zero also. And this is why you get the data landfills. Companies are always often moaning about their data landfills because it's just easier to store everything than it is to decide what you need to store and get rid of the stuff that's never going to be useful. So you just tend to store everything and you pile up all this information, which becomes- And the theory is that artificially

Speaker B:

intelligent tools, we're just chucking it all into this giant warehouse in a load of unlabeled boxes. And the assumption is that artificially intelligent tools will go in there and sort it out for us so we can find it again, essentially. Yeah.

Speaker D:

Yeah. And I think the thing that is interesting at the moment is we do seem to be in a transition. So at the moment, the volume of information that we've created, which is absolutely enormous already and only getting bigger, far outstrips our ability to comb through it and find what it is we're interested in. And so at the moment, I'm interested in these kind of forensic cases like this stamp, tracking down this stamp, and what you have to do in order to find that information. And there's a really good example is the Skripal poisoning case. And there are a couple of really interesting bits about that. One is, obviously, initially, there was all of this CCTV footage of the United Kingdom just randomly being collected. And the huge effort that was gone through to find the individuals wandering around the streets of Salisbury who they actually wanted to identify. And a lot of that was a manual-

Speaker B:

The innocent Russian tourists.

Speaker D:

The innocent Russian tourists visiting the cathedral. Exactly.

Speaker B:

123 metre high spire.

Speaker D:

Yeah. Well, quite. World famous. And well worth not visiting. Yes. So there was that effort and the fact that a lot of that was done manually. And not only that, to sort of shortcut this process, there are people known as super recognisers who are identified for their ability to be good at spotting features on people's faces. I went to school with one. Did you? Hi, Carla, if you're listening. Right. Well, there you go. You know a super recogniser. Hopefully, she'll be able to tell the timber of your voice perfectly and recognise you. So, sorry, is your point there still reliant on- Still reliant on people. But those super recognisers must be doing something that can be replicated. And in fact, facial recognition software is doing this. But it's not just about facial recognition software, because obviously they cover their faces. It's about gait recognition and all of those kinds of things. Just being able to spot someone from a single ear or a bit of a tattoo or whatever it might be. Right. But then there's a second part to this, which is that subsequently Bellingcat, the sort of investigatory kind of collective- Open source intelligence. Open source intelligence group, have been able to identify the exact GRU Colonel who has been identified as one of the people involved in the operation. And they've done that from digital forensics, effectively, looking at his old military school records and him being given an award and so on, and historical photographs, which are accessible and have now been made part of the digital record. And it suggests for things like that, it's going to be more and more difficult to commit a crime and leave no record of it, because this vast digital archive of everything from transport records, to you walking along a street, to you having been to a college somewhere, to getting an award, all of that is now a matter of public record, effectively. The difficulty is, that's great if you've got a team of hundreds of police officers and lots of people focusing on this one particular thing. But if you just happen as an individual to want to know, you know, where's that person I was at school with or something, then you're perhaps left to your own devices a bit more. And how do we help all of those individual people get the same results as somebody with a massive team at their disposal? Let's talk data. Okay. I've got some stats. Nick,

Speaker B:

tes of data, so a petabyte is:

Speaker D:

deleted, unrecoverable, right? Good, frankly. Good. I mean, you know, there's an argument for why on earth are we collecting all of this crap? What are we doing with it? There's a really good Asimov short story called The Dead Past. And it's about, it's effectively about a scientist who discovers a machine that can look into the past. And, or he discovers a better version of it. And the kind of, the regime at the time controls access to these machines and won't let anybody really have access to them other than for academic purposes. And it turns out the reason why that is controlled is because what they've realized is that instead of people looking at far off history, what society would actually do if it had unrestricted access to this stuff is spend all of its time obsessed with the very near past, i.e. more or less the present, and engaging in voyeurism on, you know, towards one another's lives. And turns out that's exactly what we're doing. We're creating all of this archive material, but we're chucking it all in the bin. We're just interested in, you know, the present is far more interesting to us than

Speaker B:

nts in Derbyshire in the year:

Speaker C:

Blue Peter episodes. That's true. We will. I've got a question. Before I do, I just want to hear from Peter. Yeah, well, so, but how do you go about deciding what's going to be of interest and what's important in the future? I mean, as we just said, so trivial things like birth and death records end up being incredibly important to historians and letters from soldiers on the front that have been collated by the army archives and things are incredibly useful set of data about what was happening any given day on any given battle and things like that. So how do you decide what characteristics can you define about information sets that helps you guess which ones you should keep and which ones you should chuck away? I think there's a, from a theory

Speaker B:

point of view, my feeling would be, well, my instinct is that it should be lots of different things. So in other words, you know, you want lots of examples of different types of information. So, because obviously information becomes much less informative at scale. You know, you're essentially sampling from the same pool over and over again. It's better to have 10 totally different data sets with a thousand records than one data set with, you know, 10,000 records by and large, broadly on average. I mean, yeah. So just to caution a bit of, you know, I suppose optimism is, you know, I mentioned that the Wayback Machine has 25 petabytes of data. Well, there's one estimate that every single book ever written, if you were to digitize it, ever written in the history of humanity would only come to just over 50 terabytes. So what that means is there's already 500 times more information in the Wayback Machine than in every book ever written. So, you know, and of course, that actually also includes every book ever written, which is in there, you know. So there is a sense in which, yes, we're losing a far, far higher percentage of information than we would have done in the past. But in absolute terms, what we're retaining is vastly more than we've ever had. You know, so there is obviously, the one is cause for concern, but the other is cause for optimism.

Speaker A:

Battle of Waterloo, right, in:

Speaker D:

be a disagreement about what happened? Yes. I think you hit on something quite interesting here, which is that it's very difficult to construct a universal truth. I mean, you would be able to probably ascertain what time the first shot was fired, something like that would be straightforward, but in terms of the, you know, there would always be ambiguities about who was in the moral right in the battle, for example. And I think, you know, there are other challenges, like the creation of faked information, whether at the time. So, you know, I mean, look at the Russians and some of the sort of artificial material they create around operations, you know, in parts of Eastern Europe and so on. That kind of information would also be chucked into the historical record. Add to that the capacity, retrospectively, to change information. So, there are financial companies now that enable you, if you buy something on, let's say you buy something on your private credit card, like a work lunch, you can go back in after the fact, when you realise, oh, that was a work lunch, that should be on expenses, and change that financial record. That capacity to change digital data, historically, exists now. I'm sure, you know, Peter will jump in in a minute about ways to prevent that from being tampered with. Spoiler alert, blockchains. But there is, I think, a large part of what the future historian will be doing is effectively trying to work out what is the real information about the event and what is faked at the time or

Speaker A:

faked subsequently. Okay. We need to draw things to a conclusion shortly. So, let's all try and

Speaker C:

round things off a bit, Peter. Well, we didn't talk about the idea of how things change in your environment sort of routinely, and you don't notice those changes. Yeah, I think that's really interesting. So, as you grow older, certain things are kind of fixed, like the colour of post boxes and buses, they don't change. But, like, next door's roof will change, and the trees will age and grow, but you don't really notice these things. And a lot of these sort of things aren't actually being measured or recorded, but may be of interest to future historians. So, you might be able to track down your neighbour's financial records 100 years hence and see that he changed his roof. But to a future economist, a centralised database of how much people spend on their houses

Speaker B:

would potentially be very interesting. Yeah, I'm kind of worried because it is hard. It's hard psychologically to notice the last time you see something. But it's also hard digitally. I mean, we can't tell when things are being deleted. You know, things that we assume are there, we can suddenly discover have gone. But yeah, I mean, like Joni Mitchell said, you don't know what you've got till it's gone. Yeah, they've paid Paradise and put up a parking lot. I think

Speaker A:

that's another podcast there, though, is this question of what's missing, and how do you

Speaker B:

identify what's missing? I think we've done it, actually. I'd like to talk about it, but I think we should save that. Because now I'm an old git, I suddenly realise what's happened to those red and white striped tents that telephone engineers used to live in. Just briefly touching upon this

Speaker A:

point, and it is kind of a bit of being a grumpy old man, is that when you talk about something that is no longer there, and maybe wouldn't it be good if we were to have that back again? People stare at you like you're a moron. And then you go, no, I'll look it up on the internet, and you can't find any record of it at all. Yeah. Okay, so look, let's stop there. Is there anything anyone

Speaker D:

wants to sort of round it off with? Just once we publish this podcast to the internet, it'll obviously be available for future generations. This wisdom is now enshrined. Yeah, don't make

Speaker C:

the same mistakes we made, future people. Until we stop paying for the service that hosts it,

Speaker B:

of course. Hit that save button now. Yeah, I mean, it is unfathomable to me. I can't imagine it, but it'd be so brilliant. Imagine if we had video from the Middle Ages. I mean, how cool would that be? To see a video of the Scottish Wars of Independence. Yeah, I mean, you know, it would be just brilliant. And of course, that's what the future generations are going to have, except for

Speaker A:

they'll all be videos of cats. Yeah. Okay. All right. We're going to stop there. Thank you, as always, for listening to the Cognitive Engineering podcast. I'm Fraser McGruer. We've been here with Peter Coghill, Chris Wragg, and Nick Hare of Aleph Insights. And until next time, goodbye.

Show artwork for Cognitive Engineering

About the Podcast

Cognitive Engineering
Welcome to the Cognitive Engineering podcast.
Welcome to the Cognitive Engineering podcast. Occasionally coherent musings of Aleph Insights. We hope you like listening to them as much as we like recording them.

About your host

Profile picture for Fraser McGruer

Fraser McGruer