>With three judges, each judging 15 maneuvers, you can do an
>analysis of variance and produce average scores for each
>judge vs each maneuver vs each pilot and determine a
>standard deviation and confidence interval for the scores.
>How well the three judge's scores track each other will give
>the standard deviation showing the randomness in scoring due
>to the judges. Three judges is plenty to get good standard
>deviations. The problem is that with so few judges the
>statistics show that the standard deviation in the scores
>will so wide that it takes huge differences to be confident
>that the scores are different.
Nonsense. Of course you can compute various statistical parameters for any data set. That doesn't mean that they are valid and indicative of anything, which they most assuredly *are not* when the data set is doesn't meet the requirements for which statistics apply. In this case the very first assumption in deriving equation in any statistics book is "for N>>1" (sample size much greater than one). Three or five is not >>1, statistical mean derived assuming that therefore do not apply. QED, right there.
But I'll continue with some of the other observations in the first page of elementary statistics books.
>This is VERY hard for people to accept, but it is a
>scientific principle just as much as the F=MA equation that
>drives the planes.
Statistics is indeed a perfectly valid mathematical field. But you are not using them correctly, as the underlying precepts of statistics do not apply to your example. For instance, comparing items that contain differences not arising from random, statistically recognized, "errors". I use the word "error" advisedly, as it's not really error, but difference.
Say you want to measure the length of a pencil. You measure it dial calipers (the same dial calipers) 100 times. You readings ranging from 5.98 to 6.02". The average is 6.0001, and the variance is .001. That's a statistical measure. So you would conclude from this that the pencil is somewhere between 6.0001" +-.095, 3-sigma. Meaning that there is a ~99.7% likelihood that it's *really* somwhere between 5.9505 and 6.0949.
Except, oops, you forgot to zero the dial first, unbeknownst to you, and in fact all the readings are biased by .2". So you have a very precise, but grossly inaccurate, measurement. Maybe you notice that an fix it. But you bought the plastic calipers from the local bargain tools bin, and the gear pitch is not really right - say, the nearest metric equivalent of the right gear pitch because that's what they had when they designed it - and all the measurements, even after correcting the erroneous offset, are subject to a *Scale factor* error. So you have an even more grossly inaccurate measurement to very high precision.
But suppose you want to, instead of finding out how long in inches a particular pencil might be, you want to find out which of several pencils is the longest. Your bias and scale factor errors are not relevant anymore. You can repeat the measurement until the precision gets small enough to be smaller than the mean difference between the closest two averages, and then you have a high confidence which is the longest. You are still grossly inaccurate about the absolute length, but which is the longest is proven to any desired degree.
Then someone else comes up with a digital caliper measuring inches, recently calibrated against a standard, and another dial caliper marked in metric. also recently calibrated.
You will get your precise but erroneous absolute answer of 6.0001, but your determination of which is longest is completely accurate.
The calibrated digital caliper gives you a equally precise determination of which is longest - probably no better or worse than yours - but the absolute length measurement is a much more accurate 5.75.
The metric caliper will also provide an equally precise measurement fo which is longest, but an absolute length reading of 146.05 millimeters.
All are equally accurate in determining the longest pencil, with statistical significance.
I would think that the analogy with stunt judging is obvious.
You method is to average a sample of 3 readings and take the variation or standard deviation of these three readings that are not only not the result of random variation, not only with an sample size violating the rules for which standard deviation is derived (N=3, N NOT>>) but also in different measurement basis.
It would be like taking the pencil readings, averageing them all together (6.0001+5.75000+146.05)/3 to get a average length of 52.6000+-80,9, 1-sigma, and therefore concluding the pencils were all the same length (and incidentally in this case, allow the possibility that some of the pencils don't actually even exist - since there's a ~67% chance that they have negative length). This is not valid statistical analysis, it's gibberish.
Same applied to stunt. Judges *absolute scores* for flights, and in fact entire career averages, *are not all the same* from one to the other. It's like comparing metric and inche measurements directly. And we are not really asking them to all come up with, say a 525 on a particular flight, and deviations from that value are *not necessarily errors*. All we should be asking them is to rank the fliers (or find the longest pencil).
I wish I had cut-and-pasted the other reply from Flitelines, since you argument is identically invalid to the last time - and thus the response it still applicable.
>Unlike judging the weight of a couple of
>rocks or the trajectory of a fly ball, people have very
>little FEEL for how statistics works. Even when you have
>trained people who have been working with this stuff for
>years, they still get it wrong. They see a chart, see the
>4.1 and 3.9 and act like that is a real difference and
>ignore the little error bar that shows a range of +-2.1.
>Happens all the time. They ignore the fact that the overlap
>in the scores is so broad there is no difference. Then
>they go and make a marketing decision and lose ten million
>dollars in wasted advertising.
Right, and if they didn't properly understand how to properly use statistics, they would do better guessing.
I am not arguing that statistics as a mathematical concept is invalid - in fact it is valid and I use it every single day (on things that are a lot more costly than tens of millions of dollars - by far- and approve analyses using *far more sophisticated* methods than you are using here although that's entirely irrelevant to this discussion). I am arguing that you are applying them in a grossly invalid way and thus drawing a wrong conclusion.
>
>Using three judges, the current 10-40 scoring, and the usual
>whole point changes that most judges use, the judges have to
>track each other PERFECTLY, exept for a 1 pt difference on
>one or two maneuvers, in order to get the variation in the
>scores small enough that a 5pt score difference means
>anything.
See above.