Showing posts with label Internet. Show all posts
Showing posts with label Internet. Show all posts

Thursday, December 01, 2011

.

Web-page mistakes

I don’t understand, sometimes, how people put together their web pages. Who really thinks that, say, pink text on a red background looks good? Seventeen different typefaces on one page? A background image that makes people’s eyes cross?

One can argue that those are all matters of taste, and, after all, à chacun, son goût. And anyway, those things are easy enough to fix: one can apply a custom style sheet right in the browser, and override all those size and color and font and background things that were specified in the web page. There are instructions and bookmarklets floating around on the web... just stick one in your browser’s bookmark bar, and then click it when you encounter a retch-inducing or simply unreadable web page.

But there are lots of web-page problems you can’t fix up in your browser, because it’s the people who put together their web pages who don’t understand, and it’s not just in matters of personal taste. Perhaps one of the most annoying of these is what I call the lazy thumbnail error.

You’ve encountered these, surely: you’ll be looking at a web site for a business or organization, and you’ll click on a page labelled The Christmas Party, or Our Staff... and the page will take forever to load. You can’t see why, though: the Our Staff page shows maybe 20 people, from the company president to the secretarial staff, each with a small photo, a name, and a short paragraph by way of a bio. No big deal here. The photos are all tiny, something on the order of 100 by 120 pixels, like my mug at the top of these pages. What’s the problem?

The problem is that the photos aren’t as small as they look, because the webmaster was lazy about creating thumbnails for the staff pics. She asked everyone to send her a snapshot, and she put them all up on the web page with HTML like this:

<img src="staffPics/jane.jpg" height=200 alt="Jane Smith">

There... that makes all the photos the same height, 200 pixels. If someone sent a larger one, it gets scaled down, nice and small. Makes for a uniform look, and the page looks great.

What the webmaster doesn’t understand is that the scaling is not done at the server, but in the user’s browser. When the browser loads the page, it sees all these IMG tags, and it requests each image URL (such as staffPics/jane.jpg, in the example above). But it has no way to tell the server that it’s only going to display it 200 pixels high, and the server has no way to know. If Jane sent a high-res portrait, eight megapixels huge, the whole thing gets sent to the browser. And then the browser has to do the scaling itself, when it renders the page.

If ten of the twenty staffies have sent large photos, that simple Our Staff page can wind up being tens or hundreds of megabytes in size, despite how tiny the headshots look in the browser. Plus, there’s a load on the browser, which has to store the full-sized images and resize them for rendering — you can sometimes see that effect when scrolling the page is sluggish.

The solution is for the webmaster to take the time to create images of the right size (or close to it) from the start. If someone sends you a 2400 x 3200 portrait, scale it down to 150 x 200 yourself, and just put that image on the web page (there are programs available for this, which make it easier to handle a lot of photos). If you want to make the larger one available for clicking, something like this will do:

<a href="staffPics/full/jane.jpg"><img src="staffPics/thumbs/jane.jpg" height=200 alt="Jane Smith"></a>

The height=200 still ensures that they’ll all be the same height, in case the thumbnails aren’t all exactly the same size (there’s no harm in letting the browser do a small amount of re-scaling). But now people won’t have to grab all those high-resolution photos unless they actually want to.

Wednesday, November 23, 2011

.

Degrees of separation

New Scientist tells us about Facebook’s analysis of the friend relationships in their social network. Only four degrees of separation, says Facebook, goes the New Scientist headline. Here’s their summary:

A few months ago, we reported that a Yahoo team planned to test the six degrees of separation theory on Facebook. Now, Facebook’s own data team has beat them to the punch, proving that most Facebook users are only separated by four degrees.

Facebook researchers pored through the records of all 721 million active users, who collectively have designated 69 billion "friendships" among them. The number of friends differs widely. Some users have designated only a single friend, probably the person who persuaded them to join Facebook. Others have accumulated thousands. The median is about 100.

To test the six degrees theory, the Facebook researchers systematically tested how many friend connections they needed to link any two users. Globally, they found a sharp peak at five hops, meaning that most pairs of Facebook users could be connected through four intermediate people also on Facebook (92 per cent). Paths were even shorter within a single country, typically involving only three other people, even in large countries such as the US.

The world, they conclude, just became a little smaller.

Well, maybe. There are a lot of things at play here, and it’s not simple. It is interesting, and it’s worth continuing to play with the data, but it’s not simple.

They’re studying a specific collection of people, who are already connected in a particular way: they use Facebook. That gives us a situation where part of the conclusion is built right into the study. To use the Kevin Bacon comparison, if we just look at movie actors, we’ll find closer connections to Mr Bacon than in the world at large. Perhaps within the community of movie actors, everyone’s within, say, four degrees of separation from Kevin Bacon. I don’t know any people in the movie industry directly, but I know people who do, so there’s two additional degrees to get to me. We can’t look at a particular community of people and generalize it to those outside that community.

There’s also a different model of friends on Facebook, compared with how acquaintance works in the real world. For some people, they’re similar, of course, but many Facebook users have lots of friends whom they don’t actually know. Sometimes they know them through Facebook or other online systems, and sometimes they don’t know them at all. Promiscuous friending might or might not be a bad thing, depending upon what one wants to use one’s Facebook identity for, but it skews studies like this, in any case.

People would play with similar things in the real-life six degrees game. Reading a book by my favourite author doesn’t count, but if I passed him on the street in New York City, does that qualify? What about if we went into the same building? If he held the door for me? If I went to his book signing, and he shook my hand and signed my copy of his book? Facebook puts a big e-wrinkle on that discussion.

But then, too, it’s clear that with blogs and tweets and social networking, we have changed the way we interconnect and interact, and we have changed how we look at being acquainted with people. I know people from the comments in these pages, and from my reading and commenting on other blogs. Yes, I definitely know them, and some to the point where I call them friends in the older, pre-social-network sense. But some I’ve never met face to face, nor talked with by voice.

So, yes, the world probably is a little smaller than it used to be. It didn’t just get that way suddenly, of course; it’s been moving in that direction for a while. Everything from telephones and airplanes to computers and the Internet have been taking us there.

Monday, July 25, 2011

.

Inventing the Internet

I’m in Québec City this week for the IETF meeting. A group of us were having dinner last evening, and at the end of the meal, as we were paying, the waitress asked us what we were all in town for. We told her were were at a meeting to work on standards for how things talk to each other on the Internet.

So she tells us about a crazy lady who comes in the restaurant every afternoon. The lady claims to have invented a bunch of things, and one thing she says is that she invented the Internet. After someone makes the required Al Gore joke, I say, well, to tell you the truth, no one at this table qualifies but we do actually have some people in our group who actually did invent the Internet. She says It’s one person who did it?, and we say no, maybe eight or ten or so... and at least four of them really are here this week.

Friday, June 03, 2011

.

Trusted identities

The U.S. Postal Service (USPS) now has a way to do a change of address online, on their web site. Nicely, it’s even all using https (SSL/TLS), keeping it encrypted, which is good.

On the first page, you select whether it’s a permanent change or a temporary one, and specify the dates.

On the second page, you select whether the change is for an individual or a whole family.

On the third page, you give the old and new addresses.

On the fourth page, you get this:

For your security, please verify your identity using a credit card or debit card. We’ll need to charge your card $1.00.
[? Help]

To prevent Fraud, we need to verify your identity by charging your card a $1.00 fee. The card’s billing address must match your current address or the address you’re moving to.

If you click the ? Help link, here’s what it tells you:

Identity Verification — Credit/Debit Card

In order to verify your identity, we process a $1 fee to your credit/debit card. The card’s billing address must match either the old or new address entered on the address entry page. This is to prevent fraudulent Change of Address requests.

Please note that the Internet Change of Address Service uses a high level of security on a secure server.

I have a few problems with this:

  1. They’re asking for credit card information in a transaction where no one expects it. They’re assuring you that it’s secure, but how does one know? This is a classic phishing tactic.
  2. They’re assuming you have a credit card to give them. Lots of people don’t have credit cards. I know some.
  3. They’re charging you a dollar to change your address online, a mechanism that’s surely cheaper for them than to have you walk into the post office to do it. That’s nuts.

To be sure, they do have to do something to make sure that people don’t change each other’s addresses as pranks, or worse. But do they really need to charge you a dollar for it? They could make a charge and then rescind it. They could give you an alternative to use a bank account, and verify it the way PayPal does, by making a withdrawal of a few cents and then depositing it back. That would also help for people who have no credit cards, but do have bank accounts — still not everyone, but it’s something.

Or you can just say, Eff this; I’m not giving the post office my credit-card information and paying them a dollar for what I can do for free, and then go into the office and waste a clerk’s time on it.

This is why there are proposals for secure identity verification. The U.S. National Institute of Standards and Technology (NIST) has an initiative called National Strategy for Trusted Identities in Cyberspace (NSTIC) that covers this sort of thing. Whether or not NSTIC is the right answer, we need to get to where we have this kind of verification available, without having to hack the credit-card system for it.

Wednesday, April 27, 2011

.

Ephemeral clouds

I’ve talked about cloud computing a number of times in these  pages. It’s a model of networking that in some ways brings us back to the monolithic data center, but in other ways makes that data center distributed, rather than central. A data cloud, an application cloud, a services cloud. An everything cloud, and, indeed, when one reads about cloud computing one sees a load of [X]aaS acronyms, the aaS part meaning as a service: Software as a Service (SaaS), Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and so on.

I use email in the cloud. I keep my blog in the cloud. I post photos in the cloud. I have my own hosted domain, and I could have my email there, my blog, there, my photos there... but who would maintain the software? I could pay my hosting service extra for that, perhaps, but, well, the cloud works for me.

It works for many small to medium businesses, as well. Companies pay for cloud-based services, and, in return, the services promise things. There are service-level agreements, just as we’ve always had, and companies that use cloud-based services get reliability and availability guarantees, security guarantees, redundancy, backups in the cloud, and so on. Their data is out there, and their data is protected.

But what happens when they want to move? Suppose there’s a better deal from another cloud service. Suppose I, as a user, want to move my photos from Flickr to Picasa, or from one of those to a new service. Suppose a company has 2.5 terabytes of stuff out there, in a complex file-system-like hierarchy, all backed up and encrypted and safe and secure... and they want to move it to another provider.

In the worst case, suppose they have to, because their current service provider is going out of business.

Recently, Google Video announced that they would take their content down, after having shut the uploads down (in favour of YouTube) some time ago. This week, Friendster announced that they would revamp their service, removing most of their data in the process.

Of course, you understand that when I say their data, here, I really mean your data, yes? Because those Google Video things were uploaded by their users, and the Friendster stuff is... well, here’s what they say:

An e-mail sent Tuesday to registered users told them to expect a new and improved Friendster site in the coming weeks. It also warned them that their existing account profile, photos, messages, blog posts and more will be deleted on May 31. A basic profile and friends list will be preserved for each user.

Now, that sort of thing can happen: when you rely on a company for services, the company might, at some point, go away, terminate the service, or whatnot. But what’s the backup plan? Where’s the migration path? In short...

...how do you save your data?

Friendster has, it seems, provided a exporter app that will let people grab their stuff before it goes away. Google Video did no such thing, and there’s a crowd-sourced effort to save the content. But in the general case, this is an issue: if your provider goes away — or becomes abusive or hostile — how easy will it be for you to get hold of what you have stored there, and to move it somewhere else?

Be sure you consider that when you make your plans.

[Just for completeness: I have copies on my own local disks of everything I’ve put online... including archives of the content of these pages. If things should go away, it might be a nuisance, but I’ll have no data loss.]

Friday, March 04, 2011

.

Reasonable network management

Back in December, the U.S. Federal Communications Commission released a Report and Order specifying new rules related to network neutrality. The rules have since been challenged in court in separate suits by Verizon and Metro PCS. They’re also under attack by the House of Representatives, though whatever they do is unlikely to pass the Senate and the president.

The Report and Order is quite long and involved, a typical federal document that runs to 194 pages (here’s a PDF of it, in case you’d like to read the whole thing). On page 135 there begins a statement by FCC Chairman Genachowski, which contains, on page 137, five points, key principles, as Mr Genachowski says, that lead to key rules designed to preserve Internet freedom and openness. That’s sort of an executive summary of the document.

I’ll note principles four and five here:

Fourth, the rules recognize that broadband providers need meaningful flexibility to manage their networks to deal with congestion, security, and other issues. And we also recognize the importance and value of business-model experimentation, such as tiered pricing. These are practical necessities, and will help promote investment in, and expansion of, high-speed broadband networks. So, for example, the order rules make clear that broadband providers can engage in reasonable network management.

Fifth, the principle of Internet openness applies to mobile broadband. There is one Internet, and it must remain an open platform, however consumers and innovators access it. And so today we are adopting, for the first time, broadly applicable rules requiring transparency for mobile broadband providers, and prohibiting them from blocking websites or blocking certain competitive applications.

In apparent response to those points, and taking transparency seriously, Verizon Wireless has recently updated their Customer Agreement (Terms and Conditions). If you scroll down to the bottom of that document, you’ll find a section called Additional Disclosures, the first paragraph of which says this:

We are implementing optimization and transcoding technologies in our network to transmit data files in a more efficient manner to allow available network capacity to benefit the greatest number of users. These techniques include caching less data, using less capacity, and sizing the video more appropriately for the device. The optimization process is agnostic to the content itself and to the website that provides it. While we invest much effort to avoid changing text, image, and video files in the compression process and while any change to the file is likely to be indiscernible, the optimization process may minimally impact the appearance of the file as displayed on your device. For a further, more detailed explanation of these techniques, please visit www.verizonwireless.com/vzwoptimization

That URL at the end lacks the http at the beginning and has not been made into a clickable link, but if you copy/paste it into your browser’s address bar, you’ll be redirected to a long page called Explanation of Optimization Deployment, full of technical details. It’s perhaps the most detailed and technical disclosure I’ve seen presented to consumers, full of terms such as Internet latency, quantization, codecs, caching, transcoding, and buffer tuning.

I have to say that the policy looks reasonable. They say that they apply their optimization (not really the right term, here, but that’s the marketing spin) to all content, including Verizon Wireless branded content. They compress images and transcode video to reach a compromise between fidelity to the original content and what’s likely to be useful on a mobile device, conserving transmission resources by doing it. But it also benefits the consumer by way of reduced data charges. They also, basically, stream the content (buffer tuning), so if you stop a video in the middle you don’t have to transmit (nor pay for the transmission of) the unwatched portion.

The only disadvantage of any of this as I see it is that there’s no way to turn it off. If you notice degradation of your video content and want to watch the original — and are willing to pay for extra data transmission that entails — you can’t.

As a first step, this looks good: it’s a reasonable policy that preserves the essence of neutrality and fits the reasonable network management model. Of course, Verizon Wireless may just be testing the water, introducing changes a little at a time, with the most benign changes first. We’ll have to see.

Tuesday, March 01, 2011

.

URL shorteners

If you’re a twit a Twitter user, you’ve likely used one or another of the URL shorteners out there. Even if you’re not, you may have run across a shortened URL. The first one I encountered, several years ago, was tinyurl.com, but there plenty of them, including bit.ly, tr.im, qoiob.com, tinyarrow.ws, tweak, and many others.

The way they work is that you go to one of them and enter a URL — say, the URL for this page you’re reading:

http://staringatemptypages.blogspot.com/2011/03/url_shorteners.html

...you click a button and get back a short link, such as this one:

http://bit.ly/eqHg3S

...that will get users to the same page. The shortened link redirects to the target page, and won’t take up too many characters in a Twitter or SMS message. It also may hide the ugliness of some horrendously long URL generated by, say, Lotus Domino.

On the other hand, it will also hide the URL that it points to. When you look at the bit.ly link above, you have no idea where it will take you. Maybe it’ll be to one of these august pages, maybe it will be to a New York Times article, maybe to a YouTube video, and maybe to a page of pornography. Click on a shortened URL at your own peril.

In addition, any URL you post, long or short, might eventually disappear (or, perhaps worse, point to content that differs from what you’d meant to link to), but if you post a load of shortened URLs to your blog or Twitter stream and then the service you used goes out of business, all your links will break at once. That didn’t used to happen, but can now. And because some of them use country-code top-level domains (.ly, .im, .tk, and .ws, for example), the services may be subject to disruption for other reasons — one imagines that the Isle of Man and Western Samoa might be stable enough, but if you’ve been watching the news lately you might be less sanguine about Lybia.

The more popular URL shorteners can also collect a lot of information about people’s usage patterns, using cookies to separate the clicks from distinct users. If they can get you to sign up and log in, they can also connect your clicks to your identity. There are definite privacy concerns with all this. URL shorteners run by bad actors can include mechanisms for infecting computers with worms and viruses before they send you on to the target site.

Of course, any URL can hide a redirect, and any URL can hide a redirect to a page you’d rather not visit. It’s just that URL shorteners are designed to hide redirects, and there are no lists of best practices for these services, along with lists of reputable shorteners that follow the best practices.

What would best practices for URL shortening services look like? Some suggestions, from others as well as from me:

  • Publish a usage policy that includes privacy disclosures and descriptions, parameters, and limitations for other items such as the ones below.
  • Provide an open interface to allow browsers to retrieve the target URLs without having to visit them. This allows browsers to display the actual target URL on mouse-over or with a mouse click. Of course, shortening-service providers might not want you to be able to snag the URL without clicking, because they may be getting business from the referrals. Services such as Facebook, while not shorteners, front-end the links posted on their sites for this reason. So we have a conflict between the interests of the users and the interests of the services.
  • Filter the URLs you redirect to, refusing to redirect to known illegal or abusive sites. Provide intermediate warning pages when the content is likely to be offensive, but not at the level of blocking.
  • Provide a working mechanism for people to report abusive targets, and respond to the reports quickly.
  • Don’t allow the target URL to be changed after the short link is created.
  • Related to the previous item, develop some mechanism to address target-page content changes. This one is trickier, because ads and other incidental content might change, while the intended content remains the same. It’s not immediately clear what to do, or whether there’s a good answer to this one.

Meanwhile, I never use URL shorteners to create links, and I try to avoid visiting links that are hidden behind them. I like to know where I’m clicking to.


Update, 11 March, this just in from BoingBoing:

Dear readers! URL shorteners’ popularity with spammers means we’ve blocked some of the big ones (at least temporarily) to cut down on the spammation. Sorry for the inconvenience! While we plan a long-term fix, just use normal URLs. You are welcome to use anchor tags in BB comments, too.

Monday, February 28, 2011

.

IP blocklists, email, and IPv6

Engineers in the Internet Engineering Task Force, in the Messaging Anti-Abuse Working Group, and elsewhere have been debating how to handle e-mail-server blocklists in an IPv6 network. Let’s take a look at the problem here.

We basically have three ways to address spam, in our goal of reducing the amount of spam in our inboxes:

  1. Prevent its being sent in the first place.
  2. Refuse to accept it when it’s presented for relay or delivery.
  3. Discard it or put it into a junk mail folder at (or after) delivery.

The last is handled by what we usually think of as spam filters, which analyze the content and other aspects of the messages. Dealing with the first involves law enforcement, as well as adoption of best practices for legal email marketers. To implement the second, we try to do various analyses during the actual transmission of the email messages, in order to respond at the protocol level with some sort of refusal. It’s rather like standing between your postal carrier and the mailbox at your house, and telling the carrier that she may put this envelope into the box, but she should take those two catalogues and the credit-card offer right back to the post office with her.

And one can actually imagine doing that, by looking at the envelopes and applying rules such as, If it’s pre-sorted, it’s probably junk, and, The more urgent it claims to be, the more likely it is to be junk. But a better way, still, would be if we could get this to happen as soon as the junk mail entered the postal system, by having a way to say, See that guy who’s dropping that pile of mail at the post office? He only sends junk, and when you see him coming just make him go away. Don’t even let him bring his pile in the door.

We have that in our email systems, in what we call IP blocklists (or blacklists). These are lists of the numeric Internet addresses of email servers that we think send so much spam that we won’t even let them come to the door. When one of these servers makes an Internet connection to one of our mail servers, we don’t even start an email protocol exchange with them — we just refuse the connection. We make them go away.

Estimates vary as to what portion of attempted spam this blocks, but at least some estimates are on the order of 90%. Despite the problems with this mechanism (legitimate mail servers do find themselves on blocklists, for various reasons, and sometimes have a hard time getting the list-managers to remove them), it’s a critical one in the fight against spam, saving a great deal of time and computing resources by cutting the spam messages off much earlier in the process.

But note that it deals with IP addresses. Today, of course, that means IPv4 addresses, those things that look like 192.168.0.1, and that there are around 4 billion of. 4 billion is a large number, but, as we’ve seen, it’s notably finite and manageable. It’s reasonable to take every IP address we ever see trying to send mail, and keep it on a list, sorting the addresses into the good ones and the bad ones. It’s feasible to block Internet connections from the ones in our list that are marked bad.

Not so when we consider IPv6. Bumping the IP address from 32 bits to 128, bumping the 4 billion up to a billion billion billion or so — the number doesn’t matter, at that point — makes it infeasible to keep a list of bad addresses. There are enough addresses there to allow the bad guys to use a new one every time, so we’d never see repeats. There are, of course, ways we can group addresses into large blocks, and know that any address we see in one of those blocks will be bad, but even that isn’t enough to make it work.

We could switch to a pass list, a whitelist of known good addresses — that would still be small enough to be manageable — and refuse anything else. But that makes it very hard for an organization to deploy a new server, or for a new organization to join in.

John Levine has one approach: leave the email system on IPv4 for the foreseeable future. Even, John points out, when many other services, customer endpoints, mobile and household devices, and the like have been — have to have been — switched to IPv6, we can still run the Internet email infrastructure on IPv4 for a long time, leaving the IP blocklists with v4 addresses, and a system that we’re already managing fine with.

Of course, some day, we’ll want to completely get rid of IPv4 on the Internet, and by then we’ll need to have figured out a replacement for the IP blocklist mechanism. But John’s right that that won’t be happening for many years yet, and he makes a good case for saying that we don’t have to worry about it.

At least not until he and I have long been retired

Monday, February 14, 2011

.

Government oversight of the Internet

Now that the protests in Egypt have led to a change in leadership — an outcome that seemed inevitable for a while, though now-former-President Mubarak denied that it would happen — I want to go back and look at a key event during the last few weeks, when the Egyptian government disconnected the country from the Internet

It appears that removing an entire country from the internet is surprisingly easy, by making changes in a system known as the border gateway protocol (BGP). This system is used by ISPs and other organisations to connect to each others’ networks, so the Egyptian government just had to order ISPs to alter the BGP routing tables to make external connections impossible.

Looking at BGP data we can confirm that according to our analysis 88 per cent of the ‘Egyptian internet’ has fallen off the internet, reports Andree Tonk of BGPmon, a site dedicated to monitoring changes in the BGP. A recent report for the OECD cited the BGP as a weak point in online infrastructure that needs to be secured — a prediction that seems to have now come true.

As the report makes clear, it’s not technically difficult, at least not for a relatively small country with a relatively centralized connection to the Internet. And we see countries such as China and Iran using similar techniques to do more selective blocking (the latter has, I understand, responded to the events in Tunisia and Egypt by joining the former in blocking access to blog sites such as this one). The issue isn’t technical, but one of policy: is the government allowed to cut off the Internet?

Of course, with countries where the government makes its own authority, the answer is always Yes. But what about in the U.S., where the government was limited, at least through the end of the 20th century, to abiding by its constitution, legislation, and a judicial system?

For one answer to that question, we can look to Senator Joe Lieberman of Connecticut, who, along with Senators Susan Collins (Maine) and Tom Carper (Delaware), introduced legislation to enhance the security and resiliency of the cyber and communications infrastructure of the United States.

The Protecting Cyberspace as a National Asset Act of 2010, S.3480 (here’s a PDF of the latest version as of this writing) was introduced last June and was entirely replaced by Senator Lieberman in December (you have to go to the bottom of page 197 of the PDF to see the new version). The December version was reported to the Senate from the Committee on Homeland Security and Governmental Affairs, which Mr Lieberman chairs (and on which his cosponsors sit). It’s now on the Senate’s legislative calendar. (The corresponding House bill is H.R.5548.)

The bill, if it should become law, would create a new operational entity within [the Department of Homeland Security]: the National Center for Cybersecurity and Communications (NCCC).

The NCCC would be led by a Senate-confirmed Director, who would regularly advise the President regarding the exercise of authorities relating to the security of federal networks. The NCCC would include the United States Computer Emergency Response Team (US-CERT), and it would lead federal operational efforts to protect public and private sector networks. The NCCC would detect, prevent, analyze, and warn of cyber threats to these networks.

The bill creates, in addition to the NCCC, quite a number of offices, councils, task forces, and programs, some of which make sense and some of which probably don’t. It creates the Office of Cyberspace Policy, whose Director is appointed by and reports to the President. It creates the Federal Information Security Taskforce, comprising executives and representatives from more than a dozen government agencies. And so on.

The entire bill is quite extensive, running well over 200 pages. And what’s frightening about it is that it puts the U.S. government right in the middle of the operation and management of the Internet within the United States and its territories — and keep in mind how central U.S. operations and U.S.-based services are to the Internet as a whole. It’s difficult to understand the effect that all this new administration will have on the operation of the Internet within the U.S., and the effect that it could have if it’s mismanaged, if it tries to respond to perceived threats, if it’s affected by right-wing zealots or other dubious elements that inhabit the U.S. political community.

I have read the bill’s summary, along with parts of the bill itself, but haven’t had time to read the whole bill yet. It’s not clear how bad it could be, nor, indeed, whether it will be bad at all... but I’m very skeptical of the result of putting such a large set of deep layers of U.S. government bureaucracy in the middle of the operation and management of the Internet. And I’m deeply worried about giving authority to make operational decisions to people who have insufficient technical knowledge to understand the ramifications of those decisions, who may have political or ideological motivations that do not coincide with what’s best for the Internet, and who can implement their decisions without the checks-and-balances oversight that protects us in other parts of our lives.

I have lots more reading to do.

Friday, February 04, 2011

.

The Internet is falling!

The big Internet tech news this week is that the last block of Internet addresses, for the version of the Internet Protocol (IP) that we mostly use (IPv4), has been allocated. Or, as the headlines are saying, we have now run out of Internet addresses. Of course, it’s filled the tech media, as above, but it’s shown up in the mainstream press as well; here it is from the New York Times, and from The Guardian.

What does it really mean, that we’ve run out of IPv4 addresses?

Well, for one thing, it doesn’t mean that we’ve run out of IPv4 addresses. The Times gets it better than the other articles, in its headline:

The Last Block of IPv4 Addresses Allocated

The last address has not been assigned, not by a long shot. IPv4 addresses are allocated to organizations in large blocks — sometimes blocks of 60,000 or so, sometimes blocks of more than 16 million. Those organizations then assign addresses within those blocks, sometimes individually and sometimes in sub-blocks. What has just happened is that the last large block of addresses has been allocated. There are still many, many IPv4 addresses available for assignment, within many of the blocks that have been allocated.

For example, IBM has a 16-million-plus block of addresses comprising all addresses that start with 9 (that is, every address of the form 9.x.x.x; they also have some of the 129.x.x.x range). Those 9.x.x.x addresses are assigned within the company’s network. Not all of them are assigned, of course; there aren’t more than 16 million devices within the company.

Similarly, Internet service providers, such as Comcast and Verizon, have large blocks of their own, some for use within the company, and some to provide to their customers.

Many companies have blocks that are much larger than they need, far more than they could ever imagine using for their normal networks. Those blocks were allocated to them in earlier times, before the worldwide web and the explosion of Internet usage, when we never thought it would matter. Or they were assigned later, when we assumed that IPv6, with many orders of magnitude more addresses, would be well deployed by now. (I’ll note that it would be very difficult, even though large portions of the allocated blocks remain unused, to reclaim the unused bits and to reallocate them.)

Let’s not be Chicken Little, here; the sky is not falling, an the Internet is not imminently doomed. Indeed, the Internet will mostly run fine, as it is, for many years yet. We’ll all be able to read our email, buy from Amazon and eBay, use Facebook, and see YouTube videos.

Eventually, we’ll be crowded out by expanding Internet use, though we have techniques to keep that at bay for a long time. What will be blocked by this are — and this should be a familiar refrain to readers here — new applications, new uses of the Internet. To move into the future, beyond email and eBay, Facebook and YouTube, we need to move to IPv6.

We have enough IPv4 addresses for now, and for a while, to accommodate putting every computer on the Internet, as long as we’re thinking of computer as we have been: desktops and laptops. Maybe iPads, too. But now add Kindles and other eBook readers. Add smart-phones. Consider that every mobile phone is a smart-phone. Do we have enough v4 addresses for all of that?

Now move into the Internet of things: add every car, because our cars need to be online. Add every television (they’ll stream video directly), every stereo receiver (streaming music, radio stations, and other audio from the Internet), every portable music player from boom box to iPod Nano. Are we getting there? Include appliances: alarm clocks, refrigerators, coffee makers. Include home- and building-automation targets: thermostats, light switches, and so on. Put in sensor networks, traffic-control and monitoring systems....

Well, given all that, we ran out of v4 addresses long ago. It’s not really the v4 address-space depletion that should be driving the move to IPv6, but the need for more address space for future applications. If you don’t think that sort of thing is important, consider this news item about electric-grid problems stemming from the recent ice storm in Texas:

FORT WORTH, Texas — A high power demand in the wake of a massive ice storm caused rolling outages for more than eight hours Wednesday across most of Texas, resulting in signal-less intersections, coffee houses with no morning java and some people stuck in elevators.

The temporary outages started about 5:30 a.m. and ended in the afternoon, but there is a strong possibility that they will be required again this evening or tomorrow, depending on how quickly the disabled generation units can be returned to service, the chief operator of Texas’ power grid said in a release.

Consider the potential consequences of intersections without traffic signals and people stuck in elevators. We’d like to shut the power down in an area selectively, killing most of it but leaving the elevators running (at least until they open on the next floor), leaving a trickle of emergency lighting, leaving the the traffic lights running. We can do that, if everything’s addressable, and the power control system is set up to allow distribution with sufficient granularity.

But if it takes a Chicken Little scare — The Internet is falling! The Internet is falling! — to get IPv6 out there, well, here it comes.

Saturday, January 29, 2011

.

Peer-to-peer means many things

I have a bone to pick with New Scientist’s headline writer. For an article about Google’s blocking of certain search terms in their instant results, the headline reads thus:

Google censors peer-to-peer search terms

In fact, the article tells us that the terms in question relate to torrent downloads. The article doesn’t mention peer-to-peer at all; it’s only in the headline. And the headline gives the wrong impression.

Peer-to-peer is not synonymous with file sharing of questionable legality. In fact, it’s not synonymous with file sharing of any kind: file sharing is one application of peer-to-peer protocols, but there’s a lot of other stuff on the Internet that works peer-to-peer.

Some examples:

  • Instant messaging
  • Networked games
  • Many voice-over-IP systems
  • Any other SIP-based application

Many of these services put two end-user systems in touch with each other, and then use peer-to-peer protocols, keeping any central servers out of the picture — and out of interference, eliminating any bottlenecks that a server might cause.

The fact that many people violate copyright by using peer-to-peer file sharing to pass around copyrighted material sometimes gives peer-to-peer a bad name. But a lot of useful stuff that’s legally solid and non-controversial is peer-to-peer as well. Let’s not tar all of those with the same brush.

Wednesday, January 12, 2011

.

Network neutrality: the battle begins

Via BoingBoing, I saw this article about T-Mobile U.K. and their new fair-use policy. It relates to the recent FCC rules that give mobile carriers a pass on network neutrality, allowing them more flexibility — we might say, allowing them to violate neutrality. While the U.S. Federal Communications Commission rules obviously don’t apply to a carrier in the United Kingdom, the tone that it sets, the tone that the Google/Verizon agreement set, is felt throughout the world.

Here’s what T-Mobile is saying in the U.K.:

From the beginning of next month, the policy will limit customers to 500MB a month, down from 1GB or 3GB, depending on the contract. If you want to download, stream and watch video clips, save that stuff for your home broadband, a document on the T-Mobile site said.

A T-Mobile spokesperson has said the new policy will apply to all customers, including those who have already signed contracts with a higher cap. A message on the company’s official Twitter account said: We have to give you reasonable notice that our fair use policy is changing.

T-Mobile is touting the change as a benefit for customers, saying they won’t be charged for going over that 500MB limit. Instead, they’ll simply be banned for the rest of the month from downloading large files or viewing video via their handsets.

Browsing means looking at websites and checking email, but not watching videos, downloading files or playing games, the company claimed. We’ve got a fair use policy, but ours means that you’ll always be able to browse the internet, it’s only when you go over the fair use amount that you won’t be able to download, stream and watch video clips.

This kind of thing is exactly what many of us fear from any rule that distinguishes wireless/mobile Internet from home broadband. Had these sorts of restrictions been in place for wireline Internet access, many innovations, many services and web sites that we take for granted now would never have been able to exist. By implication, putting such restrictions on mobile access to the Internet will block new and innovative uses and services, keeping them from ever getting off the ground.

Think about some of the stuff we’re used to, that millions of Internet users depend on every day. Oversimplifying, a bit:

  1. YouTube was enabled by the lack of limitations on data transfer. If you have to pay by the megabyte, watching videos, even low-quality, highly compressed ones, gets too expensive too fast.
  2. Facebook was enabled by the elimination of time constraints on online use. Remember when you got 50 hours a month of Internet access, and had to pay by the hour (or minute) for more?
  3. Twitter was enabled by the always on aspect of Internet access. It just wouldn’t have ever worked if when you got the urge to tweet you had to go to your computer and dial up through your modem.

The wireless carriers want to change at least some of those aspects, and if we accept their doing that we’ll accept the limitations on technology development that goes with it.

Watch YouTube at home, not on your mobile, says T-Mobile U.K.

Bollocks!, we need to say back. Change carriers while there’s still a choice, and show the other carriers what we think of that sort of policy. Even if you’ll never use more than 500 MB in a month, find a new carrier that doesn’t have this limitation. Take a stand on network neutrality before it’s too late to.

Friday, January 07, 2011

.

Downloading photos

Over time, I’ve run across a few people who have posted photos to Flickr, set the Flickr option to disable downloading, and then been dismayed to find that people were saving copies of their photos anyway. They told Flickr not to allow downloading, and people can, apparently, still download.

Some of these had set the option a long time ago. With recent changes to Flickr, they’ve made it a little clearer that this isn’t a security feature, but even with that, people don’t understand what’s going on, and how others can still download their photos.

Here’s how the Flickr setting works:

You’re looking at a photo on Flickr, and you view a specific size (click on the photo or click the Action pull-down, then select view all sizes). Above the photo in the view all sizes screen is the license information and a list of sizes, and between them it says Download and supplies a download link. Also, if you right-click (Mac: ctrl-click) on the photo, there’ll be a save image selection on the menu.

If the owner of the photo has disabled downloading, the Download line will say The owner has disabled downloading of their photos, and there will be no save image option on the pop-up menu (on some browsers it may be there, but it won’t work).

What’s important to understand is that that’s all the option does: it removes the ability to save the image (photo) using the standard browser interfaces. That is, it makes it less convenient to save the image.

But any image that’s displayed by your browser has an img tag in the HTML source, and that tag has the URL for the image. For example, if we look at this photo that I’ve posted to Flickr, and then view the HTML source for the page,[1] you’ll see the following:

<div id="allsizes-photo">
<img src="http://farm4.static.flickr.com/3136/3047554534_723e36b41f_b.jpg">
</div>

You can put that URL into your browser and get directly to the image. Of course, it doesn’t matter, because you can just right-click the image on the all-sizes page and save it. But if I had disabled downloading, when you looked at the page source you would see this:

<div id="allsizes-photo">
<div class="spaceball" style="height:768px; width: 1024px;"></div>
<img src="http://farm4.static.flickr.com/3136/3047554534_723e36b41f_b.jpg">
</div>

That extra line, the one with class="spaceball", is what blocks the right-click from being able to download the photo. But the URL is still there, and the URL still works. Anyone can still download my photo by going to the HTML source and finding the URL for it there. It would be very easy to write a Firefox add-on that would do this automatically and re-enable a save option on the pop-up menu, and it wouldn’t surprise me if someone had already written one. I haven’t looked, because I don’t really care to download everyone’s Flickr photos.

Here’s Flickr’s warning about this:

Enabling this setting also places deterrents to discourage downloading of your other sizes. (And we really do mean discourage. Please understand that if a photo can be viewed in a web browser, it can be downloaded.)

Is that sufficient? Clearly not; people are still surprised when they find that their photos are freely accessible to anyone who can get to the pages to view them. But if the web browser can retrieve the photos to show them, then they can be saved — if nothing else, they’re saved in the user’s browser cache, and a savvy user can snag them thence.

I’ve been talking about Flickr, specifically, but there’s nothing here that’s really specific to Flickr. It’s true on any web site: anything a user can view, the user can save.

There is an exception to that: there are photo sites that use Flash to show the photos. Flash is a browser plug-in that runs programs that are sent from the web server. The Flash program that displays the photos does it in a way that the browser itself is unaware of (only Flash sees it), so the browser never has the photo, nor even the URL to it. Unless someone can hack the Flash program, there’s no way the user can save the photo directly.

But even in this case, a user can capture a screen image while the photo is being displayed. The Mac’s Preview program makes it easy; use the File -> Take Screen Shot option in Preview’s menu. In Windows, pressing the Print Screen key on the keyboard will copy the screen image to the clipboard, and you can then paste it into a program such as Paint, PowerPoint, or PhotoShop. There are also plenty of other programs (I like Hypersnap, but many others are fine) that give you more flexibility.

In other words, again, anything a user can view, the user can save.

So, in general, we get back to advice that you’ve seen many times in these pages: If you want something to be private, don’t put it on the Internet.


[1] It’s easy, from the browser’s menu: in Firefox, use View -> Page Source; in Chrome, View -> Developer -> View Source; in Safari, View -> View Source; in Internet Explorer, View -> Source.

Wednesday, December 22, 2010

.

The FCC on network neutrality

Yesterday, the U.S. Federal Communications Commission (FCC) approved new rules related to network neutrality. The rules are no surprise, nor was the approval — we’ve known the basic contents for some time, and they’re right along the lines of the Google/Verizon proposal from August.

Key points of contention are these:

  1. The rules allow for paid access to higher data speeds in a way that leaves things open for non-neutral deal between carriers and providers of services or content.
  2. By stressing that carriers must not block legal web sites, the FCC leaves a gaping hole, the size of which depends upon one’s definition of legal.
  3. Wireless broadband is treated differently from wired, and the rules allow much more non-neutral behaviour on the wireless side.

On point one, let’s look at three situations, and see how they’re different. Suppose my cable provider gives me Internet service at 5 megabits per second, but offers to make that 15 megabits per second if I pay another $10/month. Is that consistent with network neutrality? Most of us would say that it is.

Now suppose my cable provider gives me Internet service at 5 megabits per second, but if I stream data from Netflix that’s capped at 1 megabit per second. Is that consistent with network neutrality? Most of us would say that it is not. If I’m paying for 5 megabits per second with no limit on the number of megabits, they should not be artificially slowing down a service they want to discourage.

OK, but let’s go in the other direction, and suppose Netflix makes a deal with my cable provider. Suppose Netflix pays them some fee, and the result is that I get my 5 megabits per second normally, but when I stream from Netflix I get 15 megabits per second. Is that consistent with network neutrality?

The new rules seem to say that it is. Is it the same as the first scenario, or different? Does it matter whether I, the subscriber pay for better service, or Netflix, the content provider, pays for better service? That’s a matter for debate. Some say it’s bad on its surface, and is inconsistent with network neutrality. Others say that as long as any provider is allowed to pay for the improved service, it’s OK (there can’t be an exclusive deal). Still others think exclusive deals are OK, as long as it’s improving service, and not penalizing someone outside the deal. Maybe, but isn’t that a relative thing?

My own opinion is that there’s a fundamental difference that hinges on who pays, who benefits, and who gets left aside. If I pay, I benefit, and all services I use benefit equally. That seems neutral. If a service provider pays, only their services benefit. That’s not neutral. On the other hand, this is rather like a manufacturer paying for preferential display of their items in a store, and we do that all the time. This is not a black-and-white issue.

On the second point, the legal web site point, we have the Justice Department’s recent action of shutting down web sites that are purported to violate copyright rules. If the FCC’s highlighting of the legal point makes it easier for Internet carriers to police the Internet, I think it’s a very bad thing. Enforcement should stay in the hands of the enforcement agencies.

To look at the third point, we need to remind ourselves that wireless, in this context, refers not to WiFi, but to broadband over cellular service. I find it hard to accept that there’s a need for or a benefit to treating the two differently. With smart phones and iPad devices, and others like them, 3G service (and, soon, 4G) is becoming as important to accessing the Internet as wired broadband is. It seems detrimental, in general, to allow — even to encourage — different levels of service through the two paths. While it may be valid to limit data rates or volumes (in general, neutrally), I can’t see the need to restrict services and applications outright, and consider the inclusion of that in the regulations to be a real problem.

That’s not to say the new regulations are all bad — the transparency requirements are good, and there are other reasonable aspects to them. But I think there are fundamental flaws in the regulations, and they should not have been issued as they are. We’re very bad at regulating technology, and bad regulations can really bite us in the ass.

Thursday, December 02, 2010

.

Domain name seizure

The Electronic Frontier Foundation notes that the U.S. government has launched a crackdown on web sites that are accused of violating copyrights. The government has done this by seizing the domain names, having the names removed from DNS resolution — the process that converts the name you give to your web browser into an actual Internet address.

Over the past few days, the U.S. Justice Department, the Department of Homeland Security and nine U.S. Attorneys’ Offices seized 82 domain names of websites they claim were engaged in the sale and distribution of counterfeit goods and illegal copyrighted works.

One major problem, as EFF reports it, is that at least some of the web sites included in the sweep are not in the business of illegal distribution, and are actually trying to do the right thing, taking down bad material when they find out about it.

What’s as disturbing, though, is the U.S. government’s attempt to censor the Internet this way. As the EFF points out, sites that are able to will only find other options, using non-U.S. DNS servers (as they are already doing). Meddling with the low-layer workings of the Internet this way is not a good thing. Shutting down web sites without due process is also not a good thing. As the Federal Trade Commission struggles with trying to address real cyber crime such as phishing and other forms of fraud, the entertainment industry has found a way to bypass the difficulties and get the Department of Justice to do preemptive copyright enforcement for them.

We criticize the government of China for blocking web sites they don’t like. It seems to me that we’re now doing the same thing. That our reasons are different matters little: this isn’t the way to deal with these issues.

This also does not make me comfortable when I think about how they might handle network-neutrality legislation.

Tuesday, October 26, 2010

.

More on Internet cafés and public networks

For my readers who aren’t terribly fond of the entries tagged technology, please stick with this one. It’s important.

Do you log into web sites from public computers, even though I advised against it four years ago? That post only scratched the surface, really: it just talked about using public computers. These days, most people have their laptops with them, and they connect them to the public wireless networks in the cafés.

Most of those networks are unencrypted. That means that you don’t have to enter a key or a password when you access the network. You just select the network name (or let your computer snag it automatically), go to a web page in your browser, and get redirected to some sort of login and/or usage-agreement screen on the network you’ve connected to. Once you click through that, you’re on the Internet.

Suppose there are twenty people in there using that particular network. All twenty of them are sending and receiving stuff through the air. How is it that I only get my stuff, and you only get yours, and we don’t see each other’s, nor the web pages of the other eighteen users? It must be that my web pages are beamed straight to me, and yours to you, right?

No. In fact, everything that everyone sends and receives is out there for all twenty computers to see. But each of our computers is given an IP address, each data packet contains the address that the packet is being sent to... and all of our well behaved computers just look at the addresses and ignore any packets that aren’t meant for them.

Computers do not have to be well behaved. Any computer in the café — or near enough to hear the wireless signals — can see everything that everyone is sending to and receiving from the network. Because the network isn’t encrypted, it’s all out there, in the clear, visible to all who care to be badly behaved.

But we aren’t completely unprotected: we have something called TLS (or SSL, depending upon the version). When the web site’s address, the URL, begins with https, your communication with that web site is encrypted and safe from eavesdropping, even if the network itself isn’t. Perhaps you don’t care who sees you reading the New York Times, but you want to be protected when you visit your bank online. Use http for the Times and https for the bank, and all is well.

And that’s important, because most web authentication just has you send your username and password openly from your browser to the web site. Anyone could snoop your ID and password as you logged in, if your connection to the web site wasn’t encrypted. But that https saves you.

But wait: I have a New York Times account, and I’ve logged into the Times web site (using https). Every time I visit the site, it knows who I am. Even when I just go to http://www.nytimes.com/ ! How does it know that, when I’m not logging in all the time?

Web sites use things called browser cookies to remember stuff about you. A cookie is a short bit of data that the web site sends and asks your browser to attach a name to and keep. Later, when you return, the web site asks if you have a cookie with a particular name, and if you do, your browser sends it. For web sites that you log into, such as your bank and the Times, the login (session) cookie is sent every time your browser touches the web site. Every time I click on another Times article, my Times session cookie is sent again. Every time I go to another page on my bank’s site, my bank’s session cookie is sent again.

My bank is set up securely, as is my credit card site, as is Gmail, as is PayPal: every contact from the login screen until I’m logged out is through https. It’s all encrypted. Not only is my password encrypted when I log in, but the session cookie that the site gives me is encrypted too, every time I send it.

The New York Times, though, doesn’t work that way: only the login itself uses https. Once it gives me the session cookie, everything switches back to http, and there’s no encryption. When I click on an article and my browser sends my cookie again, anyone in the café can grab it.

Now, the cookie doesn’t contain my password, so no one can get my password this way. But as long as I stay logged in, and the cookie is valid, anyone who has that cookie can masquerade as me. If they send my cookie to the New York Times, it will treat them as though they were me, as though they had logged in with my password.

Of course, it’s not just the New York Times that does this. Amazon does it. So do eBay, Twitter, Flickr, Picasa, Blogger, and Facebook. So do many other sites where you can buy and sell things. (All the airline sites I’ve checked do it right, using https after login.) That means that if you use Facebook while you’re at Panera, someone else can borrow your Facebook session cookie and be you, until you log out. If you stop by Starbucks and get on eBay, someone else can use your cookie to make bids from your account.

There’s some protection at some sites. Amazon, for example, will let the cookie thief browse around as you, but will want your password before placing an order... assuming you didn’t enable one-click purchasing. And depending upon the options you have set, eBay might or might not ask for your password when the thief places a bid. But Facebook and Twitter are certainly wide open, here.

To try to increase awareness of this, a guy named Eric Butler has created a Firefox add-on called Firesheep, which will make it trivial for anyone, even someone who knows nothing about the technical details of this stuff, to be a cookie thief and pretend she’s you on Facebook, or Twitter, or Blogger, or the New York Times. Eric isn’t trying to abet unethical or criminal behaviour; he’s trying to push the popular web sites, whose users will be targets of these sorts of attacks, to fix their setups and use https for everything whenever you’re logged in.

So here’s an expanded form of the warning: Don’t do private stuff on public networks, unless you’re absolutely sure your sessions are encrypted. If you don’t know how to be sure, then err on the side of caution.

Monday, October 25, 2010

.

Challenge/response still lives (barely)

Wow; I haven’t gotten one of these in a long time:

ATTENTION!

A message you recently sent to a 0Spam.com user with the subject "[redacted]" was not delivered because they are using the 0Spam.com anti-spam service. Please click the link below to confirm that this is not spam. When you confirm, this message and all future messages you send will automatically be accepted.

I wrote about challenge/response anti-spam systems about three years ago, but probably haven’t seen a challenge message in at least two years. I thought people had given up on them.

Alas, no. But if the last two years is something to judge by, they’ve at least fallen further into disfavour.

Anyway, it’s worth a re-post, then, of my three-year-old item about them. All the problems, all the reasons one shouldn’t use them, are still valid now. So, here’s the link again: head over and read (or re-read) it.

Monday, October 11, 2010

.

Search engines and their responsibility

A French court has just decided a case that will likely have a great deal of effect on online search engines if the decision is upheld after appeals. A French man had been accused of crimes relating to the corruption of a minor, ultimately resulting in a suspended sentence. He found that Google search results snagged the news items about his case, putting them at the top of search results on his name:

Given extensive press coverage of the alleged crime at the time, querying the man’s name on the popular search engine returns web pages from news publications that suggested he was a rapist, among other non-favorable descriptions.

The man argues that the statements in the online articles still available today adversely characterize him, which puts him in a disadvantageous social position when meeting new people and applying for jobs, among other situations and opportunities.

The man previously contacted Google directly to remove the defamatory articles from its search index, but the company did not do so arguing its proprietary algorithms simply return web pages in its index related to the keywords searched, that is, there is no direct human manipulation of top search results.

The result from the court was this:

The French court sided with the plaintiff, agreeing that those representations were defamatory, and ruled Google could have mitigated costs to the plaintiff by removing the pages.

The ruling ordered Google to pay €100,000, and to reimburse €5,000 in litigation costs incurred by the plaintiff. The ruling also ordered the company to disassociate the man’s name from the defamatory characterizations in Google Suggest, which suggests popular phrases while a person enters search terms in the Google search-box prior to completing a search. Additionally, for every single day the defamatory information remains in the company’s search results, Google would be fined an additional €5,000.

This decision will be disastrous for search engines and other Internet services if it stands. Moreover, it’s just horribly wrong on the surface. It makes no sense to hold indexing services responsible for the information they index, unless it can clearly be shown that they preferentially indexed certain material with a goal of creating a biased view.

Research facilities have, long before the widespread availability of Internet search tools, helped people find news items and other public information that we might rather they didn’t point to, including false information and stories that have since been debunked. We’ve always considered it the responsibility of the researcher to winnow the data.

The difference now, of course, is that the researchers are friends, neighbours, potential romantic partners, and prospective employers... and the information is much more readily available than it ever was. It’s tempting to try to make the search engines let go of obsolete information and only find the current stuff.

The problems with that idea, though, are several. It’s essentially impossible to sort out in any automated way what’s appropriate and what’s not. Even if they prefer legitimate news outlets to other sources of information, and prefer newer articles to older ones, the amount of cross-linking, re-summarizing, and background information will still show searchers plenty of nasty stuff. And who decides what the legitimate news outlets are? The search engines shouldn’t be making those filtering decisions for us.

Any mechanism that isn’t entirely automated doesn’t scale. With the untold millions upon millions of web pages that Google and other search engines have to index every day, there would be no way to respond to individual requests — or demands backed by court mandates — to unlink or otherwise remove specific information.

If this should stand, I can see that Google might have to cease operations in France. If it should spread, it might easily deprive all of us of easy searching on the Internet. That would be a far greater disaster than having a guy in Paris have to explain away unflattering news stories about a false or exaggerated accusation.

Clearing one’s name has always been a difficult challenge, and it’s only been made harder — perhaps, ultimately, impossible — on the Internet. I have a great deal of sympathy for anyone who finds himself relentlessly pursued by his past, especially when that past contains errors that weren’t his.

But this can’t be an answer to that. It just comes with too much collateral damage.

Thursday, September 30, 2010

.

Internet wiretapping

A story about impending U.S. legislation has hit the news in the last few days: Senator Patrick Leahy, along with ten co-sponsors that include Dianne Feinstein and my own senator, Chuck Schumer), has introduced S. 3804, the Combating Online Infringement and Counterfeits Act (link to PDF).

There’s a log of blog outcry about it, of course, and rightly so. I’m less worried about it than many, but I do think it’s a bad idea. Here’s why:

First, we’re meant to be a democracy, different from the totalitarian states we group together with terms such as Axis of Evil and whatnot. That means that, in general, we fit our surveillance and law enforcement into the technology, rather than limiting the technology and building it specifically to enable surveillance and law enforcement. Those who say that this is only paralleling what’s in the telephone system already are missing that the telephone system grew up from a much lower-tech starting point. Wiretaps used to be literally that: wires clipped into wired systems. And it didn’t used to be easy at all.

There’s a lot about surveillance and intelligence gathering that’s hard, and it stands to reason that those tasked with doing it should want to make it easier. Keeping it hard is actually a useful check on nascent authoritarian tendencies, and the temptation for abuse. We’ve recently had court decisions, for example, declaring it a fourth-amendment violation to use GPS tracking without a warrant. These sorts of checks are important.

There’s no saying that the sort of surveillance that S. 3804 proposes will be warrantless — and the bill does specify that a court has to approve it — but we have to remember the warrantless electronic surveillance of the Bush administration, where they bypassed no only the regular courts but also the FISA court, specifically set up to deal with monitoring terrorist action. Official abuse is a real danger.

Further, this bill doesn’t even address terrorism, nor even racketeering or other such crimes. It’s aimed at copyright infringement. Not to put too fine a point on it, but that’s a ridiculous focus for such a broad and risky remedy. There are better ways to address the problem of illegal distribution of copyrighted material, and this is an attempt to shortcut things with a blunt instrument. At least, though, it’s not as bad as the insane French HADOPI law.

Apart from official abuse, though, there’s the issue of abuse by the Bad Guys themselves, who can fool with such a system in two ways:

  1. They can take advantage of the holes themselves. Any system that allows authorized intrusion implicitly allows unauthorized intrusion as well, and we should not be so naïve as to think that won’t happen. People are corruptible, security systems are compromised all the time, and if we set it up so that any Internet communication is tappable, malefactors will make their way in and tap it.
  2. They can skirt it entirely. It will only be the normal communication channels that will have their encryption compromised, allowing officials to get the unencrypted version. If what gets put on those wires is itself encrypted beforehand — if the unencrypted version is separately encrypted — we’ve gained nothing. Once requiring specialized, high-tech, expensive machines, encryption is now easy, and any ten-year-old with a copy of PGP can do it. And anyone can create a self-signed TLS certificate to secure communication with their web site. There’s nothing the service providers can do to tap into any of that.

The result will be, as often happens with these sorts of things, that private citizens and companies that are trying to abide by the law will have their privacy and liberty compromised, while the real criminals will be able to hide as easily as they do today. If passed, this law will have some effect in the area it’s intended to... but that effect will be limited, and probably short-term.

Finally, there’s the law itself: it actually seems pretty good in its inclusion of safeguards and court involvement. There are two issues I have with it:

  1. Sec. 2324(a)(2)(B) is too vague:
    [For purposes of this section, an Internet site is dedicated to infringing activities if such site is] engaged in the activities described in subparagraph (A), and when taken together, such activities are central to the activity of the Internet site or sites accessed through a specific domain name.
    Subparagraph (A) specifies that the site must be specifically designed for these activities, be marketed for these activities, or have no significant purpose other than these activities. That provides a reasonable limitation on the Internet sites that may be targeted here. But then subparagraph (B) opens it back up in a vague way, by saying that any other site might qualify if when taken together such activities are central to the site. Subparagraph (A) clearly does not include such sites as YouTube and Facebook, but subparagraph (B) arguably could. The threat of bringing such an argument to court could exert a severely chilling effect on web sites devoted to social activities and legitimate media sharing.
  2. Sec. 2324(j) provides for a public list of sites that are alleged, without any real evidence or court involvement.
    (1) IN GENERAL- The Attorney General shall maintain a public listing of domain names that, upon information and reasonable belief, the Department of Justice determines are dedicated to infringing activities but for which the Attorney General has not filed an action under this section.
    There are mechanisms to ask to be removed from the list, and for judicial review of the case only after the Justice Department refuses the petition for removal. This amounts to an unregulated blacklist of Internet sites, and strikes me as ill advised, and possibly dangerous. There will clearly be such a list held at the Justice Department; the list should not be public. Any public list must be vetted by a court, as a necessary check on law enforcement.

I plan to write to Senator Schumer with a brief version of this post, and a pointer to the full one.

Friday, September 24, 2010

.

Notes on home networks

I have a few notes on home networks, which notes come from recent experience with some network setup issues.

Encryption: How to secure one’s network — or whether not to — continues to be a point of debate. I favour some of the arguments for being a good citizen and leaving your network open, and then making sure your computers are secure. Still, that works best if you don’t want to communicate between computers within your network, and can just wall each one off. If you do want to have them talk to each other, it’s really quite a bit of work to make sure that hackers can’t talk to them as well, and most home users will prefer to lock up the network as an extra layer of protection.

But if you’ve been following things, you’ll know that WEP (Wired Equivalent Privacy, the first encryption scheme used on wireless ethernet) is severely broken. WPA (Wi-Fi Protected Access) replaced WEP, and WPA2 replaced that. And there are personal and enterprise versions of each of those (depending upon whether it uses pre-shared keys or 802.1x authentication), and a choice of encryption algorithms (TKIP or AES). It’s dizzying for techies, so imagine a non-techie picking up a wireless router and trying to set it up.

I recently got to add new devices to two different WPA2-Personal networks that had already been working. Unfortunately, in both cases the device was a limited-function device that didn’t have a full network-configuration interface. For example, one device, with a TV interface, auto-detected the network characteristics, knew it was a WPA network, and had me enter the passphrase by nudging an on-screen letter selector. Fun.

In the end, in both cases, the new devices failed to connect. There was clearly a mismatch between the devices and the networks with respect to the WPA flavour or key. But there was no easy way to diagnose the problem, and no way to change some of the settings on the devices anyway. If they supported only WPA and not WPA2, supported only TKIP and not AES, or screwed up the algorithm to convert the passphrase into a key, I could neither determine that nor fix it.

My solution was to take the easy route: switch the network to WEP and figure that it was good enough for a low-value home network. Sigh.

Network speed: People Marketing often makes a big point of the speed of in-home wireless networks. If you’ve looked at the boxes, you’ve seen wireless routers (more accurately, switches) go from 802.11b to 802.11b/g, and now to 802.11b/g/n.

802.11 is the IEEE standard for wireless ethernet, and that’s what we network folk always used to call it (pronounced eight oh two dot eleven) before the annoying but popular term Wi-Fi was coined. The bare version, with no letters, was the first, and has long been obsolete. 802.11a and 802.11b have been in wide use for a long time, but most home routers don’t do a (which has the advantage of faster speed and less interference, and the disadvantage of covering slightly less distance). 802.11g matches (almost) the speed of 802.11a, and the recent addition of 802.11n adds a great deal of speed, along with other new features, and doubles the range.

Because of the new features of 802.11n, it’s an important step. If you have a big house to cover, its extended range is also useful. But from the point of view of speed, let’s look at what we have: 802.11b has a maximum speed of 11 megabits per second (Mb/s),[1] though it will slow down as the signal gets weaker (like, at the wrong end of the house). 802.11g (and a) goes up to 54 Mb/s. 802.11n will crank at up to 150 Mb/s (and there are proprietary extensions that push it higher, if all devices support those extensions).

That’s great if you want to move data around your house. If you back up your hard drive over the wireless network, you really want everything to support n, to get the maximum throughput for your backups. A backup of a lot of data that moseys along at, say, 5 Mb/s (a typical rate for a b router at the other end of the house) will take a long time.

But if what you need is to transfer a lot of stuff from the Internet, like for streaming movies and TV programs... well, typical cable download speeds are on the order of 5 Mb/s, at least ’round these parts. A 54 Mb/s or 150 Mb/s pipe from your router won’t help at all if you’ve only got a 5 Mb/s pipe from the Internet. The limiting factor on how fast you can stream movies (and play games and access web pages and download your email) is the slowest piece of the link — the 5 Mb/s connection to the Internet.

If you want to measure your connection speed, there are lots of sites that will help. Try My-Speedtest.com, for one.

Networking over power lines: If your wireless router won’t cover your whole house, you have devices that do wired ethernet but not wireless, or you need higher in-house speeds than you can get over wireless, there are devices such as this, which feed the network signal over the power lines in your house. You plug one of these adapters into a power outlet near your cable modem, and you connect the modem’s ethernet to it. You plug another adapter into a power outlet somewhere else in the house, and you connect it to some wired ethernet device (perhaps even a wireless router, allowing you to put the router in a more useful part of the house).

I wonder how well they work. Readers: have any of you used these? Any comments? (Here’s a silly late-night-TV-style ad on YouTube. Ya gotta love the bit with the guy holding the tangled ethernet cables.)

Of course, again, they’re marketing this as a high-bandwidth (200 Mb/s and 500 Mb/s, for the newer models) network adapter, allowing multiple HD video streams:

High Performance Powerline delivers gigabit-fast wired connection and is perfect for connecting HDTVs, Blu-ray™ players, DVRs and game consoles to your home network and the Internet.

As I said above, this won’t really do better than an 801.11g/n router at snagging HD streams from the Internet, unless you really have an Internet connection that goes at 200 Mb/s or more. As far as I know, there’s nothing remotely close to that for home use now, and won’t be for quite some time.


[1] Remember, that’s megabits per second, not megabytes; at a rate of 11 Mb/s it will take on the order of 3 or 4 seconds to transfer my MP3 of Santana’s Oye Como Va, which is about 4 megabytes.