Thursday, January 25, 2007

Copy and Paste doesn't work on the internet

Copy & paste is a crucial user interface metaphor in modern operating systems. It works both within files (for example, copying a sentence from one Word document to another), and on files themselves (for example, copying a file from one directory to another). It's always by value, rather than by reference - when I update a pasted sentence, the original doesn't automatically update too.

Copy & paste is commonly used in several different ways:

  1. Re-ordering a document, for example moving a sentence from one paragraph to another
  2. Creating lists of similar items, for example copying a spreadsheet formula then editing the copy slightly
  3. Managing versions, for example copying a document then archiving the copy
  4. Re-using content, for example copying an image into a presentation
  5. Creating personal copies of documents, for example copying a file from a USB drive
There are problems in using copy and paste for cases 4 and 5. For example, in case 4, the BBC website reuses the BBC logo on every page. But this isn't done with copy and paste - otherwise, they'd need to store it thousands of times, and they'd have to find and update every single copy whenever the logo changed. Instead, each web page copies it by reference (using its URL), rather than by value.

Similarly, corporate email systems use copy and paste for case 5. Consequently, if I send a 1MB attachment to 50 people, it uses up 50MB of storage. When the master version changes, everyone still has the old copies.

And frankly, copy and paste is a poor tool for version management (case 3), especially in the enterprise. Rather than relying on file names and saving copies, a version management tool can show proper audit history and tracking.

Copy & Paste is very personal - it's rooted in the old paradigm where everyone worked on their own PCs (with their own copies of files) and communicated every so often, rather than sharing content as a community from the very start.

The internet is different, because of the URL. Rather than emailing that 1MB attachment to 50 people, if you put it on a website it will only use up 1MB of storage. And when it's updated, people will automatically see the most recent version. That's because the URL works by reference, not by value.

This requires a shift in mindset. For example, in a connected world, what is the point of having more than one copy of any song? There should be a single URL - for example, www.beatles.com/lovely_rita.mp3, and everyone in the world should access this URL every time they want to listen to the song (you can imagine aggregation sites cataloging and presenting songs together). Except for temporary cacheing, there would be no other copies - browsers would not allow it, unless explicitly stated. Your phone, stereo, TV and PC would all access this URL, with the same user profile to control permissions.

Without copy & paste, a solution to Digital Rights Management is possible. If someone wants to copy and edit a song, for example to make a cover version, they should have to apply to get copy permissions.

If you look back at the list above, only 1 and 2 - content creation, rather than re-use - are appropriate uses for copy & paste on the internet. Copy & paste don't work well on the internet, because of the power of the URL.

Tuesday, January 23, 2007

Citizens own their data, not the government

For the last couple of years, there have been arguments in the UK about how much personal data government departments should hold, and how they should manage it. These arguments increased recently when Tony Blair announced that existing databases, holding police records, social security files, tax information, and pensions data, would be integrated.

From Tony Blair’s perspective, this will lower costs and enable new uses for existing data. His critics say that this is yet another invasion of our privacy with potentially dangerous consequences – what if a hacker or a bad government got their hands on the information?

Both sides are correct. It makes no sense to cripple government departments with overlapping databases that don’t link to each other, especially if this means citizens having to repeat the same information to each department. But no one is convinced that there are adequate safeguards.

What’s worse, no one knows the scale of the issue. Who can tell me exactly which pieces of information the government knows about me, which departments have access to it, and what they are using it for? Until I’m told, how can I (or indeed a senior civil servant) sensibly pass judgment as to how and when the databases should be joined up?

It’s my data (and yours) in these databases – the government is holding it on our behalf. So I think I have a right to know exactly what’s in these databases about me, who has access to them, and what they’re using it for. There should be a secure website where I could conveniently check my details from all government departments, and update them as necessary.

This right of data ownership would provide the safeguards necessary for me to trust the government to join databases up.

Any time the government asks for more data, it would be immediately obvious, and it would have to be justified. Existing information should be used instead where possible. This avoids the current ‘creep’, where more and more information is held – this week the police DNA database, next week satellite car tracking – without adequate debate.

By putting citizens in control of their own information, and showing how it is used, data quality is likely to go up (always an issue with databases), and trust in government will increase (as suspicions about data usage are replaced with facts).

There’s even an opportunity for government to grant data access to external organizations, too (with my approval). For example, when I fill out a magazine subscription, there could be a button saying “get address from UK government”. Then whenever I move house, the magazine subscription is automatically updated based on UK government data. And on the government website, I could see all the external organizations that can look up my address. Providing a standard ‘user profile’ would give a massive boost to the internet economy.

The government should continue to press for efficiency and integration in its IT systems. But it should ensure that all citizens have the right to see all their data, edit it where appropriate, and find out exactly who else can see it and what they use it for.

Thursday, January 18, 2007

Connecting Digital Devices

The biggest problem in the consumer electronics industry is reportedly how to get devices to work together.

Everyone wants to sell the 'central device' that coordinates all others in the house, but no one can agree on what it will be - the PC, the set-top box, the games console, or even a new 'home server'.

What's more, no one can agree on what it will do - centrally store your photos, music and video, allow appropriate access to iPods and other computers, administer fridges, ovens and other objects, or just enable content to be shared across devices.

None of these ideas sound particularly enticing to me. Why would I want a complicated machine that manages all others in my house - can't they manage themselves? So long as I can listen to the same music on my iPod and stereo, or use the same address book on my phone and my PC, then I'm ok. If I really wanted to turn my oven on remotely, I would want to do it via any device - phone, PC, console - not just the central one.

What’s needed is not a central device, but a way for each device to publish content and services for the others to consume.

The solution is maddeningly obvious! We already have the technology - it's the humble URL! Why not give each device a URL, and why shouldn't each device publish its content to the web, for example via RSS feeds?

For example, my phone should offer its address book, call history, photos, and music via authenticated RSS (or Atom) feeds. And my PC should be subscribed to this RSS feed, keeping it automatically synchronized.

The beauty of this model is that it’s straightforward – people are used to typing URLs, and there are plenty of browsers around to enable it.

And it’s flexible. For example, rather than installing a web server on your phone, Google or Yahoo or Vodafone could host its content for you, by asking you to download synchronization software onto your phone. Then, you could see your phone’s photos and address book in Gmail or Yahoo! Mail. And from there, you could access them from any other device.

It’s also manageable. My web hosting company manages data backups, rather than me fiddling around on a PC or set-top box. If a device breaks or gets stolen, then I still have the data and I can turn services off, or delete content, remotely. If my train goes under a tunnel, then I use the latest RSS data cache for my email, rather than a live link.

Of course, subscription can work both ways. My phone could also be subscribed to my Gmail contacts RSS feed, so rather than fiddling around with phone keypads I could type my address book in Gmail, and it would synchronize automatically. Either Microsoft’s SSE extensions to RSS, or the Atom Publishing Protocol, could handle this.

I don’t think the power of the URL has sunk in, especially in the consumer electronics industry. The REST approach should be drilled in to product designers, and in particular the use of RSS / Atom for subscriptions and synchronization.

The whole problem of connecting digital devices boils down to publishing and syndication, and the solution to this problem is to use the internet.