Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Thursday, January 12, 2012

.

No mean feat

Physicist Stephen Hawking turned 70 last weekend, and has been living with ALS — amyotrophic lateral sclerosis — for nearly 50 years. Usually, the disease is diagnosed in patients over 50, and they die within a few years. I was reading an article in Scientific American about Dr Hawking’s longevity. The article contains an edited interview with Dr Leo McCluskey, an ALS expert at the University of Pennsylvania.

One answer, in particular, struck me:

Sci Am: What has Stephen Hawking’s case shown about the disease?

Dr McCluskey: One thing that is highlighted by this man’s course is that this is an incredibly variable disorder in many ways. On average people live two to three years after diagnosis. But that means that half the people live longer, and there are people who live for a long, long time.

The mathematician in me rose up at that: no, on average does not mean that half the samples are on each side of the average. Average refers to the arithmetic mean — take a bunch of numbers, add them, and divide by the count (how many numbers you added) — and it’s easy to show, by example, how that’s wrong.

Suppose we had five patients with ALS. Suppose four of those patients lived for one year following diagnosis, and one lived for eleven years. 1 + 1 + 1 + 1 + 11 = 15, and 15 / 5 = 3. So on average, people in this sample lived for three years... and only one of the five (20%) survived more than even one year. Given Dr Hawking’s experience of on the order of 50 years, he could offset about 25 patients who succumbed after one year, and still give us a three-year average.

The problem with the arithmetic mean is that it’s easily skewed by outliers. In the extreme example here, if 96% of the samples are 1 and 4% are 50, we get an average of 3 — three times the normal value. That means that with such a situation, the average is useless in giving us any reasonable prediction of what to expect. More generally, if the numbers are widely variable, the average doesn’t tell us anything useful. If we have nine patients who made it through 1, 2, 3, 4, 5, 6, 7, 8, and 9 years, respectively, what do we tell the tenth patient who shows up? 5 years, on average, sure, but, really, we might as well tell him to take a wild guess.

Averages are useful when the values tend to cluster around the arithmetic mean, particularly when the number of samples is large. They’re also helpful in analyzing trends, when we look at the change in the average over time... but, again, we have to be careful that a new outlier hasn’t skewed the average. Sometimes we adjust averages to try to compensate for the outliers — for example, we might eliminate the top and bottom 5% of the samples before taking the average.

Another common error is to confuse the mean with the median. The latter is often used in financial reporting: median income, median purchase price for houses, and so on. The median is a completely different animal from the mean. It’s, quite simply, the middle value. List all the sample values in increasing order, and pick the one in the middle (or one of the two in the middle, if the number of values is even).

In the first example above,[1] if we write the values as 1, 1, 1, 1, 11, the median is the value in bold: 1. In the second example, we have 1, 2, 3, 4, 5, 6, 7, 8, 9, for a median of 5. You can see that in the first case, the median is not related to the mean, while in the second case it’s the same as the mean. It’s also the case that the mean (or average) is an artificial value that might not appear in the samples, whereas the median is, by definition, one of the sample values.

Also by definition, at least half the sample values are greater than or equal to the median (and at least half are less than or equal to it). In other words, Dr McCluskey’s statement would have been true (at least close enough) had he been talking about the median survival period, rather than the average. Medians are also less susceptible to skewing by outliers, as you can see from the first example.

But as the second example shows, when the numbers are all over the place, neither is of much use in predicting anything.


[1] My examples use small numbers of values for convenience. In reality, both mean and median require a fairly large sample size to be useful at all.

Thursday, December 15, 2011

.

Patterns in randomness: the Bob Dylan edition

The human brain is very good — quite excellent, really — at finding patterns. We delight in puzzles that involve pattern recognition... consider word-search puzzles, the Where’s Waldo stuff, and the game Set. We’re also great at giving patterns amusing interpretations, as we do when we fancy that clouds look like ducks or castles — or when we claim to see images of Jesus in Irish hillsides, pieces of wood, paper towels, and store receipts. Remember the cheese sandwich with the Virgin Mary on it, which sold on eBay for $28,000 in 2004? Miraculous, indeed.

It’s with the knowledge that we find apparent patterns in randomness that I approach this puzzling aspect of the random play feature of my car stereo. I’ve stuck in a microSD card that has about 4000 songs on it. I’ve put it on random play. And it appears to be playing songs in random order.

But it sure seems to be playing a lot of Dylan.

Bob, not Thomas. I like Bob Dylan, of course; that’s why I have quite a bit of him on the microSD card. But, for instance, on one set of local errands, it played two Dylan songs, something else, another Dylan, two other songs, then another Dylan. Four out of seven? Seems a bit odd.

Now, I know that if you ask a typical person which sequence is more likely to come up in a lottery drawing, 1-2-3-4-5, or 57-12-31-46-9, he will say not only that the latter is more likely, but that if the former came up he’d be sure something was amiss. In fact, they’re equally likely, and are as likely as any other pre-determined five-number sequence, but the one that looks like a pattern is one we think can’t be random. Similarly, it’s certainly possible to randomly pick four Dylan songs out of seven — or even four in a row, for that matter. And if there’s a bug in the algorithm that the audio system uses, why would it opt for Dylan, and not, say, Eric Clapton or the Beatles, both of which I also have plenty of on the chip?

So I played around with some numbers. Let’s make some simplifying assumptions, just to test the general question. Assume I have 20 songs from each artist, and a total of 4000 songs (and, so, 200 artists). If I play seven songs, how likely is it that two will be by the same artist?

It’s easier to figure out how likely it is that there won’t be repetitions. The first song can be anything. The likelihood that the second will be of a different artist than the first is (4000-20)/3999, about 99.5%. The likelihood that the third will differ from both of those is (4000-40)/3998. Repeat that four more times and multiply the probabilities: there’s a 90.4% chance of seven different artists in seven songs... meaning that there’s about a 9.6% chance of at least one repetition. Probably more likely than we might think.

Let’s look at Dylan, specifically. I have about 120 of his songs on there (3% of the total; maybe I should delete some, but that’s a separate question). What are the chances of having no Dylan in seven songs? No Dylan for the first is 3880/4000, 97% (makes sense: 3% chance of Dylan in any one selection). Continuing, no Dylan, still, for the second is 3879/3999. Repeat five more times and multiply: 71.3% chance of no Dylan, so there’s a 28.7% chance of at least one Dylan song if we play seven.

What about the chances of at least two Bob Dylan songs... a repetition of Dylan? Well, we figured out no Dylan above. Let’s figure out exactly one, and then add them. For the first to be Dylan and none of the others, we have 120/4000 * 3880/3999 * 3879/3998 * 3878/3997 * 3877/3996 * 3876/3995 * 3875/3994. About 2.5%. It’s the same for one Dylan in any other position — the numerators and denominators can be mixed about. So the chances of exactly one Dylan song out of seven is 2.5 * 7, or 17.5%. Add that to the chances of zero, 71.3 + 17.5 = 88.8%, so there’s an 11.2% chance of at least two Dylan songs in a mix of seven songs.

In other words, it’s a better than one in four chance that I’ll hear at least one Bob Dylan song, and a better than one in ten chance that I’ll hear at least two of them every time I take a 20- or 30-minute ride. Thrown in some confirmation bias, where I forget about the trips that had Clapton and the Beatles and Billy Joel and Carole King, but no Dylan, and I guess the system is working the way it’s supposed to.

But, damn, it plays a lot of Bob Dylan!

Friday, October 30, 2009

.

Gimme an "F"!...

Many of you have probably heard about the Governator’s playing with steganography — specifically, an acrostic:Governor Schwarzenegger’s veto letter

Of course, everyone’s buzzing about whether it could possibly have been accidental, as his staff claims, or whether it just had to be intentional. My opinion: I think it was intentional. But, opinions aside, let’s take a look at it for a moment, without trying to pull out actual numbers (there are plenty of folks posting statistical analyses that do have numbers, so you can look for those if you like).

The chances, of course, of writing some text that spans seven lines and happens to have the first letters of each line spell “Fuck You”, are exceedingly, vanishingly, teeny-tiny. Really, really, really, really small. Even more so if you consider capitals, and the paragraph break between the two words, forming the “space”. Minuscule.

But not zero. It is possible, however unlikely it may be, for it to happen by chance.

But back up and consider that there are many combinations that might have appeared there and been thought noteworthy. It could have said “kiss ass”, for example, or “eat shit”. It didn’t even have to take all seven letters: if it said “bite me”, and the first or seventh letter were unrelated, we’d probably still hear about it and wonder if he’d meant to do it. We’d even wonder about “x no way x” (substitute your favourite irrelevant letters for the x’s).

So the chances of getting something that we’d notice and put in the newspaper are, while still not at all high, not as tiny as it seems when we look at the single phrase that did show up.

And, in fact, the chances that there’d be something sensible there at all, even if it were, say, “tadpole” are actually pretty good (especially if you consider incomplete strings, like the “no way” example). This is what sucks the bible-code lunatics in, when they think they can find secret messages from God hidden in the bible. On the surface, it seems to us that “discovering” a convoluted pattern that we can make sense out of in light of something we know (or would like “evidence” of) means something. In fact, it means only that we’re good at artificially finding meaningless patterns.

How likely is it that “1 2 3 4 5” will come up in tonight’s lottery drawing, in that order? How likely is “37 12 83 7 22”? Exactly the same (both very, very unlikely). Yet we perceive the former to be “impossible” — if it should ever show up, we could be sure of a fraud investigation, and a likely re-draw — while accepting the latter as a normal “random” set of lottery balls.

If we should fairly and thoroughly shuffle an unarranged deck of cards, what’s the probability that we’d end up with the four aces on top? What’s the probability that the top four cards, in order, will be the jack of hearts, the three of spades, the seven of hearts, and the ten of diamonds? The former is actually much more likely than the latter (because the order doesn’t matter... 24 times more likely, 1 in 270,725 vs 1 in 6,497,400), but we’re inclined to think otherwise. We do not have a good intuition about probabilities, and we overemphasize what appear to us to be “obvious” patterns.

Of course, there’s a difference between playing with the layout of bible text until we can find some pattern that we interpret to be a fuzzy message about the JFK assassination... and straight out seeing “fuck you” as an acrostic in something someone wrote yesterday. Yeah, I think Governor Schwarzenegger (or his staff) was having some fun. And my sincere thanks to the Governor for making today’s blog topic easy.

Friday, September 18, 2009

.

Public misunderstanding of studies

Over at Bioephemera, Jessica Palmer agree with Language Log’s Mark Liberman in his admonition against the use of “generic plurals” in science reporting. Language Log:

This would lead us to avoid statements like “men are happier than women”, or “boys don’t respond to sounds as rapidly as do girls”, or “Asians have a more collectivist mentality than Europeans do"” — or “the brains of violent criminals are physically and functionally different from the rest of us”. At least, we should avoid this way of talking about the results of scientific investigations.

The reason? Most members of the general public don’t understand statistical-distribution talk, and instead tend to interpret such statements as expressing general (and essential) properties of the groups involved. This is especially true when the statements express the conclusions of an apparently authoritative scientific study, rather than merely someone’s personal opinion, which is easy to discount.

The problem, in case you don’t see it from what’s quoted above, is this (I’m going to make some details up, just to give an example):

Suppose some researchers do a study in which they ask people how happy they are, on a scale of 1 to 10. Suppose that they ask 50 men and 50 women, and the average happiness rating for the men is 7.3, while the average score for the women is 7.1. Now suppose that the study is reported in the news with the statement that “men are happier than women.”

Or let’s be even more straightforward: suppose the 50 men and 50 women are simply asked, “On the whole, are you happy?” 37 of the men and 36 of the women say, “Yes.” And the newspapers report that, according to a recent study, “men are happier than women.”

Of course, George reads that over his morning coffee, and says, “Hey, Martha. It says here that I’m happier than you. Ha! I always knew there was something wrong. Maybe you need some of that Prozac stuff.”

But we can’t generalize a finding based on average aspects of a group... to particular individuals in the general population. Martha may be far happier than George, and the study doesn’t say otherwise. George just doesn’t understand.

Of course, the problem isn’t limited to generic plurals with no statistics behind them. We could report that a study shows that “men are 50% more likely than women to get into traffic accidents,” but that wouldn’t mean that I am 50% more likely, just because I’m a man. There are other reasons, which the study might or might not go into, that are the causes of the difference, and the study just shows one correlation.

So it’s important to word these reports in a way that doesn’t invite that sort of misinterpretation. It’s important for a number of reasons:

  • The media already often get the details wrong in reporting scientific studies. It makes it worse to compound that with confusing reporting.
  • The media often highlight the wrong bits, in efforts to get catchy headlines and “interesting” copy.
  • Readers don’t understand statistics, and misinterpretation is likely even when the stats are there. Don’t make it worse by eliminating them.
  • Readers are prone to generalize results beyond what’s valid, and they’ll likely apply a group trent to specific individuals, as in the example above.
  • Readers don’t understand the limitations of studies. Reporters should try to talk about one or two key limitations.
The first two are nicely demonstrated by the British newspaper The Telegraph. Back in June, they reported on work done by a student, Sophia Shaw, at the University of Leicester. The preliminary findings, according to Ms Shaw: “We can see from the results that sexually experienced men are more likely to coerce women in sexual situations; even more so if they believe the women to be sexually experienced.” But the Telegraph reported (the article has since been removed from their web site after the criticism of it, but you can read discussion of it) that the work “found that the skimpier the dress and the more outgoing the woman, the less likely a man was to take no for an answer.”

In The Telegraph’s competition, The Guardian, Ben Goldacre seemed to enjoy tearing the former’s report apart:

Women who drink alcohol, wear short skirts and are outgoing are more likely to be raped? “This is completely inaccurate,” Shaw said. “We found no difference whatsoever. The alcohol thing is also completely wrong: if anything, we found that men reported they were willing to go further with women who are completely sober.”

We often say that the public needs to be better educated with respect to science and critical thinking. This is a good place to start... and the news media need to be among the educators.

Friday, July 17, 2009

.

Gun control doesn’t work?

Today’s guest blogger, while I’m paying attention to presentations at CEAS, is frequent commenter Ray.

Oh, wait... lookee here (on page 51 of the associated report), where we find that for the year 2008/2009[1] the number of murders by gun in the U.K.[2] was a whopping 38, down from 53 in the previous year.

I expect it’s just a coincidence that the U.K. has strict laws concerning gun ownership.

Let’s see, since the population of the U.S. is around five times that of the UK, that number is equivalent to 190 gun-related murders in the U.S. Hmmm... that number doesn’t jibe too well with the typical reported annual U.S. rate of around 10,000.

I expect it’s just a coincidence that the U.K. has strict laws concerning gun ownership.

 
Ray, thanks for filling in with a topic so dear to my heart!


[1] According to the report, “estimates are derived from interviews carried out between April 2008 and March 2009 (BCS year ending March 2009).”

[2] It actually covers only part of the U.K., and doesn’t include Scotland and Northern Ireland. That doesn’t change Ray’s point very much, though: it changes the approximate factor from 5 to 5.6, which changes 190 to about 210. The order of magnitude is the same: we still have about 50 times the per-capita gun-murder rate here as there.

Saturday, July 12, 2008

.

That's a mean median

I was struck by something in the lede of a recent NY Times article, headlined “Minimum Wage Increases Faster Than Median Wage”:

In the last few years, the minimum wage in New York State has increased almost 40 percent, while the average pay for hourly workers has risen much more slowly, not even keeping pace with inflation, according to a report released Thursday by the federal Department of Labor.

The median wage paid to the 4.1 million hourly workers in the state was $12.03 last year, meaning that more than two million New Yorkers earned less than that, the report from the Bureau of Labor Statistics showed. That was about equal to the median national hourly wage of $11.95 — about $25,000 a year for a 40-hour work week.

The headline and the rest of the article say “median”, but the lede says “average”.

Another word for “average” is “mean”, and they are not the same as “median”. To get the average, you sum all the values and divide by the number of values. To get the median, you list all the values in ascending order, and take the middle one. One major difference is that a few very high (or very low) values will skew the average (mean), but will not affect the median. Reports on house values and salaries usually use the median for that reason.

To take an extreme example, suppose the values are these:

1, 1, 2, 2, 2, 2, 690

The median value is 2, in bold, (the one in the middle of the list of seven items), but the average is 100 (the list sums to 700, then divide by 7). So if those represent salary levels (say, take-home pay in thousands of dollars a month), it’s clear that the median gives a better view of the real situation than the average does.

Of course, it’s possible for the median to produce a strange result too. Another extreme example:

1, 1, 2, 3, 100, 800, 893, 900, 900

Here, the median value is 100 and the average is 400. In this case, both give a rather weird view of the data, but the average is probably more meaningful than the median, if either can be said to have much meaning at all. (And that’s why there are things in statistics like deviation and skew.)

But people often don’t understand the difference between the two terms, and incorrectly use them interchangeably. And for most people, it doesn’t matter: they get the idea, and don’t care about the statistical details.

The New York Times should get it right, though.

[I did point this out to the reporter, who told me that he knows the difference, and that the error in the lede was introduced by an editor, who changed it before it went up.]

Friday, May 23, 2008

.

What grade is an “F”?

In Good Math, Bad Math, Mark Chu-Carroll talks about a USA Today article about grading systems in schools. The article focuses on arguments about what an “F” means, and the “fairness” (or not) of systems that assign “A” to scores of 90-100, “B” to 80-89, “C” to 70-79, “D” to 60-69, and... “F” to 0-59. How can it be right, those who would change the system say, to have 10 points each assigned to the other grades, and a whopping 60 points assigned to “F”? [Update to clarify: the quotes below come from the USA Today article, not from Mark's blog.]

In most math problems, zero would never be confused with 50, but a handful of schools nationwide have set off an emotional academic debate by giving minimum scores of 50 for students who fail.

Officials in schools from Las Vegas to Dallas to Port Byron, N.Y., have proposed or implemented versions of such a policy, with varying results.

Their argument: Other letter grades — A, B, C and D — are broken down in increments of 10 from 60 to 100, but there is a 59-point spread between D and F, a gap that can often make it mathematically impossible for some failing students to ever catch up.

Mathematically impossible? Say what?

Mark addresses the issues nicely, pointing out that there might be valid arguments for a change, depending upon exactly what they’re trying to do — the basic problem is in converting from percentage scores to letter grades and back. There are a couple of things I want to say beyond what Mark said.

I want to highlight a discussion in the comments section of Mark’s post: for most of the sorts of exams we’re talking about, there’s some sort of “baseline” score, a score that one could expect to get by random chance, even if one knew absolutely nothing about the subject. On a true/false test, 50% is that baseline, a grade that could be achieved by a pre-school child who randomly filled in answers.[1] On a multiple-choice test with four choices for each question, the baseline is 25%.

So there’s a perfectly reasonable argument that the baseline score, the score that could be expected of an entirely ignorant student, should be where the “F” grades start. You should have prove you know something in order to get even a “D”.

And, now, what’s that about “mathematically impossible”? Here’s some clarification:

“It’s a classic mathematical dilemma: that the students have a six times greater chance of getting an F,” says Douglas Reeves, founder of The Leadership and Learning Center, a Colorado-based educational think tank who has written on the topic. “The statistical tweak of saying the F is now 50 instead of zero is a tiny part of how we can have better grading practices to encourage student performance.”
And that statement, itself, is a glaringly good example of a classic mathematical dilemma. This is not a statistical problem, and trying to apply statistics to it is silly. Students do not have “a six times greater chance of getting an F,” unless they are playing the random-chance game with the test, trying to beat the baseline score with guessing and odds alone (and even then, it’s not six times). And changing the grading system is not a “statistical tweak”.

The tests are measuring (or trying to) whether you know the material. What Mr Reeves is saying is that your knowledge combines with some roll-of-the-dice chance to create your score, and that even if you learn the material, you have the odds against you because of the scoring bias. What I’m saying is that knowing the material, not random chance, is what you need to pass the test.

Then, too, is the question of fairness. Excuse me: exams are not meant to be fair, in the sense of distributing grades equally, or of giving students who haven’t learned the material an equal chance to those who have. Someone who can correctly multiply two three-digit numbers 50% of the time is clearly better off than someone who can only succeed 10% of the time, or not at all. Yet we still might consider 50% to be failure at that exercise.

Suppose we decide that to measure performance for the football team tryouts, we should divide students’ performance in doing push-ups in groups of 10. If you can do 90 or more, you get an “A”. 80 to 89 is a “B”; 70 to 79, a “C”, and 60 to 69, a “D”. If you can’t do 60, you get an “F”. We’re grading it that way because we’re using it for qualification for the football team: if you don’t fail, you can be on the team. Would anyone think it “unfair” that 60% of the data points denote failure?

Of course not. We would say that we require that our football players be able to do at least 60 push-ups. We would say that we have minimum standards. We wouldn’t say that every kid should have an equal chance to be on the team — only that each should have an equal chance to try out. Same with academics. We have minimum standards, and it’s not unfair to say that when you don’t meet that minimum, you fail.

There is the issue that a zero (or a score close to it) on one test, when put into the overall grade formula with the rest of the semester’s work, will be hard to overcome. That’s true, but it would be a unusually hard-nosed teacher who would look at one grade of, say, 10, mixed with other grades of 65 or 70 or so, along with demonstrated effort... and wouldn’t allow some sort of make-up work to replace the inordinately low score.

We don’t need to revamp the grading system. We need to have students who take the work and the learning seriously, and teachers who look at the whole picture and act (and grade) accordingly.

We have a tendency to want to stuff everything into a formula, to come out with a well defined answer. Life is not like that.
 


[1] In fact, it would be as hard to get a zero on a true/false test as it would be to get a 100. I’d almost be inclined to give an “A” to someone who got none of the answers right, on the assumption that it was done as a joke, because you’d have to know the material quite well in order to manage to do that badly.

Tuesday, April 08, 2008

.

Oh, Monty, Monty, Monty...

Ah, the Monty Hall Problem. It’s been done to death on the Internet, and long before — back in the pre-Internet days when networking was often by dial-up, mouse-clicking hadn’t been invented yet, and people posted pontifications by typing green text on black 80-character-wide screens. That scene at the beginning of 2001: A Space Odyssey, where the apes are going nuts? Yeah, the movie doesn’t say so, but the obelisk is one of the doors and it’s all an argument about the Monty Hall Problem.

But the Internet has a way of making things long settled resurface, resurrect, as it were, like zombies in a George Romero film. Everything you could possibly want to know about the Monty Hall Problem (well, except for the endless, endless, endless arguments about it) is summarized nicely on Wikipedia. Or you could go to Google for lots more, including this cute video that explains the problem and its solution quite well.

So what can the New York Times — or anyone else — possibly add? Well, the Times reports that Dr Keith Chen, of Yale University, has identified flaws in behavioral research going back at least 50 years... flaws rooted in the Monty Hall Problem:

The economist, M. Keith Chen, has challenged research into cognitive dissonance, including the 1956 experiment that first identified a remarkable ability of people to rationalize their choices. Dr. Chen says that choice rationalization could still turn out to be a real phenomenon, but he maintains that there’s a fatal flaw in the classic 1956 experiment and hundreds of similar ones. He says researchers have fallen for a version of what mathematicians call the Monty Hall Problem, in honor of the host of the old television show, “Let’s Make a Deal.”

The Times article then goes on to explain the flaw in an experiment involving monkeys selecting preferred colours of M&M candies. OK, look: as someone who sorts M&Ms by colour and saves the best for last (orange, of course), I find the surprise only to be that it took them 50 years to figure out that they got it wrong.

But, well, it’s not really all about M&Ms, and Dr Chen contends that they have, indeed, been misinterpreting studies involving choice for all these years:

Even when the experimenters use more elaborate methods of measuring preferences — like asking a subject to rate items on a scale before choosing between two similarly-ranked items — Dr. Chen says the results are still suspect because researchers haven’t recognized that the choice during the experiment changes the odds. (For more of Dr. Chen’s explanation, see TierneyLab.)

“I don’t know that there’s clean evidence that merely being asked to choose between two objects will make you devalue what you didn’t choose,” Dr. Chen says. “I wouldn’t be completely surprised if this effect exists, but I’ve never seen it measured correctly. The whole literature suffers from this basic problem of acting as if Monty’s choice means nothing.”

In any case, you should check out the article, if only to play the cool simulation that they have there.

 

I chose my apparel, wore a beer barrel
And they rolled me to the very first row
I held a big sign that said “Kiss me I’m a baker,
and Monty I sure need the dough!”
Then I grabbed that sucker by the throat until he called on me
’Cause my whole world lies waiting behind door number three

—— Jimmy Buffett

Sunday, March 16, 2008

.

Home field advantage?

I’ve never understood the idea that the “home team” has any “advantage” in a sporting event, for being at its home field, or court, or stadium, or arena. One might say that that’s because I’m not much of a “sports” kind of guy, and so I don’t know these things, but it’s really never made sense to me. I play volleyball now, and have played other sports in the past, and I know that I play the best I can regardless of where I’m playing.

People have told me that the advantage come from the support of the fans, or because they “know the field” better, but can that really be it? One would think that major sporting venues would be pretty much standardized, at least as to the playing field, and that any differences would be known by all teams over time. And wouldn’t any pump-up that can be attributed to cheering fans of one team be counterbalanced by extra “adrenaline” in the other team, in a desire to give an “Up yours!” to the home team and their fans?

Well, wouldn’t you figure: a couple of guys in Germany have actually studied it, and written a paper (summary in New Scientist, abstract, full paper (PDF)).

Some excerpts from the paper, which is generally chock full of mathematics:

In typical soccer reports one can read that a team is particularly strong at home (or away) or that it is particularly successful in scoring goals (or has a particularly good defense) and that it is just playing a positive series (Lauf in German) or a negative series. Here we show that the actual data do not support all of these pieces of common knowledge of a soccer fan.

[...]

It will turn out that there is indeed no additional team-specific home fitness. In contrast, the concept of the goal fitness can be backed up by the data, but only as a minor effect.

[...]

Whenever a team plays better at home than expected (in terms of its ∆GH - ∆GA) this effect can be fully explained in terms of the natural statistical fluctuations, inherent in soccer matches.

And then this, from the concluding discussion:
Probably, for a typical soccer fan also this statistical analysis will not change the belief that, e.g., his/her support will give the team the necessary impetus to the next goal and finally to a specific home fitness. Thus, there may exist a natural, maybe even fortunate, tendency to ignore some objective facts about professional soccer. We hope, however, that the present analysis may be of relevance to those who like to see the systematic patterns behind a sports like soccer. Naturally, all concepts discussed in this work can be extended to different types of sports.

Well, there we go. My guess about the lore of the “home field advantage” is that it derives from confirmation bias: the fans tend to remember the home wins, and forget the home losses.

Wednesday, September 12, 2007

.

Life expectancy

The statistic we call “life expectancy” is one of the most misunderstood and, therefore, one of the most misused. In an article about life tenure of Supreme Court justices, the New York Times makes one of the standard mistakes:

Life tenure today, of course, has a dimension that would surprise the Constitution’s framers; since 1900, the average life expectancy, now 77 years, has increased by 30 years.

Before we look at this, we have to note that when we talk about “life expectancy”, we usually mean what’s properly called life expectancy at birth — we can, in fact, statistically calculate life expectancy at any age: how much longer you can, on the average, expect to live after having already survived to the given age. Infant mortality and death from early-childhood diseases are significant factors in reducing the life expectancy at birth.

So here’s what’s true: the U.S. life expectancy at birth was, indeed, about 47 years in 1900, and is about 77 years now. The Times is right about that. But assuming that that tells us that the Constitution’s framers expected Supreme Court justices to live, on average, to be 47 or less — the life expectancy at birth in the late 18th century was even lower — is silly. Consider that Thomas Jefferson was 83 when he died, Benjamin Franklin was 84, James Madison was 85, and John Adams lived to be 90. Those ripe old framers knew how things went.

The problem with looking at it the way the Times does is that in 1900, as in the 18th century, a great many infants and young children died. Reaching one’s 5th birthday was no small accomplishment then, but, having done that, you could expect to continue for a good many years, most likely long past 47.

Life expectancy in ancient RomeFor a good visualization of this effect, look at the graph on the right, which shows the life expectancy at various ages in ancient Rome (click the graph to get to the web page that it comes from). The life expectancy at birth was 25 years, but a child who lived to the age of five could expect to reach 48, and someone who became a senator at 50 had, on average, 17 years of life and service still ahead of him.

Another interesting web site to try this out with is this Austrian one, which lets you plug in your date of birth and gives your life expectancy as of now. According to that site’s program (which I haven’t verified), we can see this:

Birth year    Life expectancy    
at birth
   
Expected age at death
as of now
198081.7684.15
196075.7982.61
194067.1883.92
192056.2491.94

While the life expectancy at birth goes steadily down as we go back in time, the adult life expectancy doesn’t change much. In fact, someone born in 1920, having made it to the age of 87, still has about five years left, at least in Austria.

Since Supreme Court justices are generally appointed in their 40s or 50s, they can certainly expect reasonably long service on the court, and that was also the case with the Jay and Elsworth courts at the beginning. Some of the very first Supreme Court justices lived to be 78, 84, 87, and even 92, so one can hardly say that our founders expected them to die relatively young.

Anyway, the Times article is otherwise interesting, discussing a proposal, popular among judicial scholars, for changing the tenure of the justices.

Tuesday, January 30, 2007

.

Counting the homeless

New York City is in the process of taking a count of its homeless population. Of course, that's not as easy as taking a count of homeful people, where you just stop at each house. For the homeless (or maybe the feds would prefer, now, to call them "people with very low roof security") one has to go to the places where they hang out, where they sleep, where they find meals... and with the homeless one doesn't always know where those places are.

There's been criticism from various advocacy groups that the city's counts seriously understate the problem — that they're very inaccurate, and result in numbers that are way too low. The city, in its turn, says that these are just estimates and it doesn't really matter if they're too low. That's as may be, but when you think about it you realize that there's little value to a count whose accuracy is that uncertain.

I heard on the radio yesterday afternoon that in an attempt to make the count more accurate, the city will be placing “decoys” for the city's counters to count — people who will appear to be homeless, but who are not.

WTF?

That was my first thought on hearing that. But then Mayor Bloomberg explained: they can use the count of the decoys to adjust the count of the homeless. If they know the percentage of decoys that were missed, they can scale the count of the true homeless accordingly, to compensate for mis-counting.

An interesting idea (with an unfortunate name, likening the homeless to ducks being hunted, but never mind that). OK, let me think about it some more:

  1. In order for this to work at all, the decoys must be invisible with respect to the counters — that is, the counters can't know which are the decoys and which are the real homeless. Otherwise, the presence of the decoys will skew the count.
  2. It seems that the counters have to be pretty much invisible to the homeless. A good portion of the homeless population would otherwise hide, suspicious of or frightened by the counters. The counters have to be low-key, and can't go accost all the homeless people.
  3. Number 2 means that the decoys can't reliably know whether or not they've been counted, to report that fact later. Yet in order for this to do anything, the city needs to have an accurate count of how many decoys were counted and how many were missed. I don't see how they can do that with any accuracy.
  4. This mechanism, by its design, can fix accidental errors for situations where the counting parameters are known. “Oops, no one went down 53rd St,” can be corrected for. The critics are not concerned about these sorts of errors; they're worried that the counters simply don't check certain places because they aren't aware that the homeless congregate there, or they're concerned for their own safety. In those cases there won't be decoys there either, and this mechanism will have no way to compensate for those situations.
  5. Expanding on number 4, this mechanism's accuracy is fundamentally related to the extent to which the answer is already known. The proportion of decoys in a given area must approximate the proportion of real homeless in that area in order for the scaling to do any good. If you're trying to count Hasidic Jews and you send out a load of “decoys” wearing dark suits and hats, you'd better send lots of them to Bensonhurst and only a few to Greenwich Village. If you do it the other way around, your decoys will way overstate the error in the Village and understate it in Bensonhurst. It's like that: if you don't already know what areas the homeless tend to be in, you don't know where to send the decoys.

So here's my thought after hearing the explanation and doing a little analysis of it:

WTF?