Stuka Stunt Control Line Forum
Archive, 2000–2021 · recovered from the Internet Archive
Forums › Stuka Stunt Main Forum

Statistically Studying Scores

Stuka Stunt Main Forum · 117 of 117 known posts recovered

godzilla · Oct 14, 2004 05:39 PM

edited#0 source
I have been in my Six Sigma training all this week. The instructor for this module has a Masters in statistics (all of these guys do). Aside from being a Master Black Belt he is very cool guy. I mentioned to him that I flew CLPA and he thought it was very interesting. He was also interested in my side class project.

There have been a few discussions on judging and how difficult it is to "grade" a particular judge's performance. It certainly is difficult, as there are only a few methods available. My concern was being able to find a "coaster" in a group. That would be a judge who is either clueless and faking it or simply using a "bracket" to score a particular flier. In my mind all good judges should roughly track each other, coasters would simply stick to a range. In other words, no one flies a perfect pattern manuever to manuever, no matter how good they are. Some manuevers are better than others. A good judge will see a good manuever and go up, and conversely see a bad manuever and go down. If you had a bunch of good judges, no matter the overall score for each manuever, the cylce should be universal. Up on good, down on bad.

Let me first say that I understand now what Brett has been talking about, there are no *extremely simple* tests (like just eyeballing the scores, even if you had all of them) that will show one judge different from another. However, using statistical software it was really pretty easy.

I have the fake data that I used for the tests. It shows 4 judges and 5 fliers. 3 are expert and 2 advanced (I could add more). Two judges are bogus in two completely different ways. The scores are in the same range, and the totals are comparable. In fact, without using some particularly powerful means, the answer is in no way obvious.

I have this spreadsheet available. I would be interested to see if anyone else could solve the "puzzle" and find the two bogus scores. If anyone wants the spreadsheet e-mail me, or I can put it on the UHP website for download.

Also, I am interested to know what methods have been used, or are being used to rate the performance of judges at Nats and Team Trials. I am sure that it has come up historically, but I have never been privy to it.

Keep in mind, this is all fake. It is also just for fun and something for me to do to practice using my new software. Don't anyone go nuts here.

Brett Buck · Oct 14, 2004 07:59 PM

edited#1 source
>Let me first say that I understand now what Brett has been
>talking about, there are no *extremely simple* tests (like
>just eyeballing the scores, even if you had all of them)
>that will show one judge different from another. However,
>using statistical software it was really pretty easy.

Perhaps. What parameter do you think corrresponds to a random number? What distribution does this random parameter? What is your sample size?

I'm not trying to bait you into anything - it's my contention that there are no parameters that correspond to random numbers, and the sample size is far too small, meaning the assumptions underlying the derivations of the parameters does not apply.

In your example, you of course found what you were looking for. But how do you know that your "coaster" was not just finding all the maneuvers of a similar quality, vice "bracket scoring". I don't think you can determine this statistically.

Brett

klelmore · Oct 14, 2004 11:32 PM

edited#2 source
You probably can discover a judge whose scores have a different variance structure from other judges is the data set is big enough. While this one doesn't sound large enough for a Pearson-like approach, there may be some resampling methods that would help and something along the lines of an ANOVA approach (Fisher would be so happy). I tend to prefer non-parametric approaches whenever possible and take the power hit...

Scores are numerically strange beasts since they aren't normally distributed: they're constrained to a finite interval.

I'm intersted because my day job is as a research meteorolgist at the National Severe Storms Lab, in Norman, OK, and statistics is one of my primary resreach tools.

I'll bite; send me the data (I'll ingest it into S-Plus). It might be useful when I teach my up-coming resampling stats course this spring (I'm adjuct faculty and have to teach every other year -- for free, no less).

Send it to me at work:

kim DOT elmore AT noaa DOT gov

Cheers!

Kim ELmore

Ted Fancher · Oct 15, 2004 12:22 AM

#3 source
'zilla,

At the risk of sounding like a broken record, I think I have to align my untutored opinion with Brett's intellectually superior one.

The fact that you have supplied the numbers to be assessed (as I understand it, not based on actual observation of the maneuver to be quantified; if not I apologize and retract my observation)implies that you have selected numbers which would stand out to a statistical analyst; i.e. two "judges" give a 29 and 30 to a wingover and the other two give it a 16 and a 39. Statistically, assuming that they were all supposedly objective measures of a "given" phenomenom, the two extremes would be immediately suspect.

Brett's point, as I see it is that since the observations are in no measurable way "objective" any one of the ( for instance) four could, in fact, represent the "legitimate" value of the maneuver.

As I read the problem, the only way a "statistically" relevant assessment could be made would be if, in fact, there was a means whereby the wingover could be objectively assessed as meriting a "specific" number representing its "quality of perfection" relative to an observable objective standard. Thus, each judge's score could be compared to the predetermined "objective" assessment and the relative merits of each judge's score thus derived.

As I understand your premise, you are "assuming" that the majority of judges will assess the maneuver "accurately" which will thus tell you the other judge/judges are faulty in their assessment for whatever reason.(without reference to the maneuver itself by those evaluating the judge's performance)

Again, the problem as I see it is the assumption that there is a "correct" value for a maneuver which can somehow be determined by the examination of scores subjectively provided by judges, none of whom have omniscience in terms of their assessment.

Yuuch, that sucked didn't it? Did my assessment make any sense at all?

Ted

godzilla · Oct 15, 2004 10:09 AM

#4 source

>The fact that you have supplied the numbers to be assessed
>(as I understand it, not based on actual observation of the
>maneuver to be quantified; if not I apologize and retract my
>observation)implies that you have selected numbers which
>would stand out to a statistical analyst; i.e. two "judges"
>give a 29 and 30 to a wingover and the other two give it a
>16 and a 39.

I did not devise a test to show how smart I was. Just the opposite, I had no idea if I would see a difference at all. In fact, I was really frustrated that the scores looked perfectly acceptable and reasonable. It took a very specific test to even show any difference.

Statistically, assuming that they were all
>supposedly objective measures of a "given" phenomenom, the
>two extremes would be immediately suspect.

You would think so wouldn't you!

If it were obvious, I would not have bothered to post. In fact, the scores are all similar, but not the same. The means are similar, the distributions are normal, and all of the data sets overlap. Perfectly benign. In fact, I ran a cooralation check and the test was invalid because it would not find the bad judges. It showed a P value of 0, which means that the sets are related, when in fact they are not.

BTW I used 4 judges and 5 fliers for 2 rounds. That is 15 manuever scores per flight per judge. That is 600 data points. I think that is PLENTY (about 5 times more than needed to be a good test). The fact is, that it takes very few flights to study the trends if you are looking at the variance from manuever to manuever.

No offence, but you guys are not giving me much credit here. Maybe you should look at the test before you comment. You assume a lot. Brett appears to want me to tell him the method by which the scores were derived. Well, shoot, that would give away the puzzle!

What methods are being used to study scores now?

Ted Fancher · Oct 15, 2004 11:32 AM

#8 source


>
>No offence, but you guys are not giving me much credit here.
> Maybe you should look at the test before you comment. You
>assume a lot. Brett appears to want me to tell him the
>method by which the scores were derived. Well, shoot, that
>would give away the puzzle!
>
>What methods are being used to study scores now?

'zilla,

I apolgize for commenting. Won't happen again.

Ted

godzilla · Oct 16, 2004 12:37 AM

#23 source
>>I apolgize for commenting. Won't happen again.
>
>Ted

OK. I appreciate your understanding. Thanks.

Congrats on winning the Vice Prez!

Brett Buck · Oct 18, 2004 12:55 AM

#53 source

>BTW I used 4 judges and 5 fliers for 2 rounds. That is 15
>manuever scores per flight per judge. That is 600 data
>points. I think that is PLENTY (about 5 times more than
>needed to be a good test).


Well, that could possibly be true if the goal was to determine the average score for all maneuvers from all fliers. But of course that's not what you are trying to do. You are trying to determine which scores from each of 4 judges are "anomalous". That means your sample size is 4 - not 600.

>The fact is, that it takes very
>few flights to study the trends if you are looking at the
>variance from manuever to manuever.

Which has absolutely no relevance. A low variance from maneuver to maneuver could simply mean that that judge happened to find the maneuvers of similar overall quality just as easily as it represents "bracket scoring". Or that one of the judges put, say, disproprotionate weight on one element that was lacking throughout, and the others did not. For example, if we had a wise-guy combat flier who automatically downgraded all square maneuvers by -25 points because he "knew" that they were exceeding the specified corner radius (although he in fact is basing that opinion on things not directly associated with the flight path of the airplane, but on previous analysis).


Brett

klelmore · Oct 15, 2004 11:04 AM

#7 source
There are two ways to approach the problem. It sounds like your approach compares the judged scores to some "truth," but in this case we have no truth by which to judge the quality of the judge's scores.

From what I've seen in the posts, about the only thing that can be analyzed is an intercomparison between judges. In a multivariate sense, we (ideally) will see one set of scores attributed to one judge that has a different variance (or whatever dispersion measure is used) structure from the other judges. That judge would be decalred an "outlier" (not necessarily an "out-liar") and would be deemed as creating scores based on a different, unknown process. All we know is that judge is *different*, not necessarily wrong.

If an analyst suspects something about the processes involved in creating the scores, a much more (statistically) powerful test can be generated to test for the difference between two different processes, or a known process and a different, but unknown process. This might be what's being done here.

I'm a bit concerned about the nomality assumptions: the scores are distributed over a finite interval, whereas normally distrubtued scores will range from minus to plus infinity. That said, this assumption is *always* violated, and the key is to use a statistical method that is robust in the face of such violations. Some methods (like the F-test) are very sensitive to these violations while other methods (for example, the t-test) are not. We can also always transform the data so that it fits into a parametric distribution with which we can work easily. How well all of this works depends on the test employed and whether or not we are looking for differences between two processes, one of which we know about.

I'm curious: aside from academic curiosity, why has this come up? Is this an issue within the stunt community?

Kim Elmore

Brett Buck · Oct 17, 2004 10:56 PM

#50 source

>I'm curious: aside from academic curiosity, why has this
>come up? Is this an issue within the stunt community?


No.

Brett

ferocious · Oct 15, 2004 01:21 PM

edited#13 source
>The fact that you have supplied the numbers to be assessed
>(as I understand it, not based on actual observation of the
>maneuver to be quantified; if not I apologize and retract my
>observation)implies that you have selected numbers which
>would stand out to a statistical analyst; i.e. two "judges"
>give a 29 and 30 to a wingover and the other two give it a
>16 and a 39. Statistically, assuming that they were all
>supposedly objective measures of a "given" phenomenom, the
>two extremes would be immediately suspect.

No, Brad's experiment was to see if he could use statistics to find out if some judges were actually giving better scores to better maneuvers or if they were just scoring randomly within a certain range. So what is being tested is how well the judges scores track each other from maneuver to maneuver.

>
>Brett's point, as I see it is that since the observations
>are in no measurable way "objective" any one of the ( for
>instance) four could, in fact, represent the "legitimate"
>value of the maneuver.

The objective measure is the average of the scores given by the judges. To evaluate the judges you look at how well their scores track each other, maneuver to maneuver, and how much they vary from the average. They have to agree on which maneuvers were better or worse and they have to have a fairly small variation from the average, or the scores don't mean anything at all.

>As I understand your premise, you are "assuming" that the
>majority of judges will assess the maneuver "accurately"
>which will thus tell you the other judge/judges are faulty
>in their assessment for whatever reason.(without reference
>to the maneuver itself by those evaluating the judge's
>performance)
>
>Again, the problem as I see it is the assumption that there
>is a "correct" value for a maneuver which can somehow be
>determined by the examination of scores subjectively
>provided by judges, none of whom have omniscience in terms
>of their assessment.
>
>Yuuch, that sucked didn't it? Did my assessment make any
>sense at all?
>
>Ted

No, it didn't. If there is no "correct" value for a maneuver, then judging them is pointless. Until someone comes up with a computerized image analysis/scoring system, judges will do the job.

We ask the judges to assess the maneuver according to certain rules and come up with a numerical value. We average the judges scores to get a number that is closer to the correct value than any individual judge's score may be. The more judgements, the closer to the correct value. Maybe Kim can shed more light on the problem.

Brad's question is entirely different- how to figure out if the judges are doing a good job. A judge COULD judge expert simply by selecting random numbers from 30-36.(Brad's question) If he did, how can you tell that he is not actually judging? You could never tell just by looking at the scoresheets for a few flights. All the judges would probably be scoring in the 30-35 range at first glance. Checking how well the judges scores track, maneuver to maneuver, could show it up though.

klelmore · Oct 17, 2004 12:38 AM

edited#36 source
OK, I'll set down the handle and be a statistician for a bit. There may be some confusion about what's going on here *statistically*, so I'll put on my professor's hat (the silly, flat thing) and have a go...

>We ask the judges to assess the maneuver according to
>certain rules and come up with a numerical value. We
>average the judges scores to get a number that is closer to
>the correct value than any individual judge's score may be.
>The more judgements, the closer to the correct value. Maybe
>Kim can shed more light on the problem.

The ONLY reason anyone ever takes the mean of scores, which are also known as "data," is because the analyst suspects the data contain random effects (noise, which is by definition indistinguishable from random). Were this not the case, only one judge (data point) would be used. No one does this, of course, because no one trusts that one data point accurately depicts what's being measured. This is why we have statistics: to try to get a better handle on what's being observed.

An important digression follows...

Any time a mean, or avereage, is used, there are assumptions about the underlaying distribution the data come from. One is that the data come from random samples from a normal (Gaussian or "bell-shaped") distribution with a specific mean and variance, both of which are unknown to the analyst. THIS DOESN'T MEAN THE VALUES THEMSLEVES ARE RANDOM! It means that we choose the values from this distribution randomly. From the data we extract in our sample, we try to discover the unknown mean (and sometimes variance) of the Gaussian distribution.

Another assumption is that our data result from a random sample of the population. That is, we have available all possible judges' scores, and we randomly pick N of them. In truth, neither of these assumptions are ever satisfied. So what? depemdimg on the statistical techniques we employ, violating of the underlying assumptions can pead to worthless, or misleading, results. One way od handling some of these violations is with resampling statistics, but not even this is a panacea.

Let's cut to the (statistical) chase: in a statistical sense, we do silly, and utterly unsupportable things with our scoring system. To get a good mean, we'd need at least 5 data points. So, to get a decent mean, 5 judges is the absolute *minimum.* Ten would be much better.

Then, we need to look at whether the difference in scores is statistically significant at some level (let's use 95% or a p-value of 0.05). Assessing statistical significance addresses the following question: are scores of, say, 480 and 465 *statistically different*? With only three judges, the answer is an unequivocal *NO*. With ten judges, the answer is "maybe," depending on how much variance there is within the scores. One way to reduce the variance is to exclude the high and low scores and take the mean of what's left. This is called a "trimmed mean," and it's used all the time.

However we choose to proceed, we then compute confidence inetrvals about the mean, and use *these* to assess if scores are statistically different. Graphically, if the mean value of one score is outside the confidence interval of another score, the two scores are declared "different." (If you're still reading at this point, this idea comes from a famous 20th century statistician named Tukey.)

But, we don't do this. Why? Judges are expensive beasts to acquire (not in $ but in experience). We'd want equally, and very, experienced judges. Can you imagine rounding up ten good judges for each active circle at every contest? Even then, let's be honest: if someone told *you* that your score of 480 was really no different than the other guy's score of 465, would you simply shrug and say "OK?" Me neither. As people, we simply don't work that way. Hence, the system we use.

>Brad's question is entirely different- how to figure out if
>the judges are doing a good job. A judge COULD judge expert
>simply by selecting random numbers from 30-36.(Brad's
>question) If he did, how can you tell that he is not
>actually judging? You could never tell just by looking at
>the scoresheets for a few flights. All the judges would
>probably be scoring in the 30-35 range at first glance.
>Checking how well the judges scores track, maneuver to
>maneuver, could show it up though.

This is the crux of the matter. I absolutely assure everyone here that, with enough data (this is an essential caveat), a posterior analysis can quite clearly find judges whose (unknown) criteria for judging specific maneuvers differs. In addition, we can also easily find judges that generate random numbers within a bracket. However, we cannot do this based on any single analysis of a single flight, or (worse yet) a single mneuver of a single flight.

Unlike what someone else suggested, we need *NOT* know beforehand why a judge scored as they did to discover this, but we do need a track record for judges. Once we discover systematic differences (and with enough data, we will), we then inquire about how the judges generate their scores. As someone else suggested, each judge may be using somwhat different, and perfectly defensible, criteria. This is a perennial problem with human scoring -- we certainly aren't special in this regard.

For discovering different scoring criteria, the bigger the differences between judging criteria, the less data we need to discover the difference and vice versa.

Everything I've discussed here is all quite straightforward, linear statistics.

Aside from finding slacker judges (I'd be surprised if there were many at all, though I've never competed), philosophically what would we do with such data, if we had them? Do we really want automata judging our flights, or do we want well-trained people, who provide us their comments, judging us? I argue for the latter, with an emphasis on "well-trained."

Note that we can even do statistics based on judges comments! If we did, which would we discourage: judges who reward hard corners (complete with bobbles), or the judges who prefer smooth, round corners? The possibilities are endless -- as varied as personalities and individual tastes.

Sorry about the length, but if discuss the statistical anaylsis of scores and judges (or anything else), we need to do so properly. The comment about statistics and lies is only appropriate when statistics are misused, which is easy to do if there is an underlying agenda. I have no such agenda. If we're gonna play stats with scores, let's do it right.

Kim Elmore


ferocious · Oct 15, 2004 12:47 PM

#10 source
>
> In your example, you of course found what you were
>looking for. But how do you know that your "coaster" was not
>just finding all the maneuvers of a similar quality, vice
>"bracket scoring". I don't think you can determine this
>statistically.
>
> Brett

The judges are supposed to be using the same criteria for judging. If they are, their scores from maneuver to maneuver should at least go in the same direction- up is better, down is worse. If one judge's scores do not "follow the crowd" and track with the other judges, that judge cannot be using the same scoring criteria. If he views the maneuvers as all "of similar quality" and three other judges thought the squares were noticeably better than the round loops there is a big problem with the judge's training.

Brett Buck · Oct 15, 2004 02:03 PM

#15 source
>>
>> In your example, you of course found what you were
>>looking for. But how do you know that your "coaster" was not
>>just finding all the maneuvers of a similar quality, vice
>>"bracket scoring". I don't think you can determine this
>>statistically.
>>
>> Brett
>
>The judges are supposed to be using the same criteria for
>judging. If they are, their scores from maneuver to
>maneuver should at least go in the same direction- up is
>better, down is worse. If one judge's scores do not "follow
>the crowd" and track with the other judges, that judge
>cannot be using the same scoring criteria.

Or he's using the same scoring criteria with different weighting on the different elements. Or he using he same scoring criteria with less overall "gain" on the score. Or he's seeing something more correctly than the other judges. All of which are perfectly legitimate ways to score the contest, and none of which represent errors in the scoring, and certainly not random errors following any recognized distribution for which statistical measures have been defined.

And certainly, using N-1 judges as the "standard" by which the "correct" value is determined (and used to determine the Nth score is "incorrect") is completly bogus, if N=3 or 5. If you have three judges, it's a very resonable to posit that two of them are bracket scoring, and the third may be doing it "righter". But the "right" guy would be the standout and would get tossed.

> If he views the
>maneuvers as all "of similar quality" and three other judges
>thought the squares were noticeably better than the round
>loops there is a big problem with the judge's training.

Which is not solved or even detected accurately by statistical analysis.

Brett

ferocious · Oct 16, 2004 09:17 PM

#30 source
> Or he's using the same scoring criteria with different
>weighting on the different elements. Or he using he same
>scoring criteria with less overall "gain" on the score. Or
>he's seeing something more correctly than the other judges.
>All of which are perfectly legitimate ways to score the
>contest, and none of which represent errors in the scoring,
>and certainly not random errors following any recognized
>distribution for which statistical measures have been
>defined.

If you go by the rule book, the judges are all supposed to be doing the same task. Using different weights on the elements is wrong, even though it is done(judges routinely ignore the 5 ft radius and reward smoothness). If a judge is using a narrower range(less "gain")his scores will track up and down with the goodness of the maneuvers and there is no problem. If one judge is doing a better job, it will show up in the variance between judges and you might be able to show the other guys how to do it better.


DMoon · Oct 16, 2004 09:37 PM

#33 source
(judges routinely
>ignore the 5 ft radius and reward smoothness).


uhhhh...not around here. Hard with a bobble gets more than medium and smooth.

NO ONE HAS EVER ANSWERED THAT QUESTION!!!!

Which is a worse?

Can you?

ferocious · Oct 18, 2004 04:39 PM

#75 source
undoubtedly any kind of a wiggle or wobble during a sharp turn is worse than missing the radius. It's much easier to see than the turn radius and hence much easier to score. Otherwise, why would the pilots work so hard to eliminate bobbles and not work equally hard to get shraper radii??

Brett Buck · Oct 18, 2004 12:13 AM

#51 source

>If you go by the rule book, the judges are all supposed to
>be doing the same task. Using different weights on the
>elements is wrong, even though it is done(judges routinely
>ignore the 5 ft radius and reward smoothness).


Where does it say that? How to weight various things is not specified, and you have to have weights for different elements to have useful scores. The event is *subjective* at it's root. If you don't like that, stay the heck out of the event.

And you're simply full of it if you think the corner radius is not considered. It's just that it is *only one of many* elements to be considered. Or do you propose that we DQ everybody who exceeds 6 ft radii?

This obsession with 5 ft is a classic symptom of someone who doesn't know very much about the event.

> If a judge
>is using a narrower range(less "gain")his scores will track
>up and down with the goodness of the maneuvers and there is
>no problem. If one judge is doing a better job, it will
>show up in the variance between judges and you might be able
>to show the other guys how to do it better.

Sort of right. The purpose is not to come up with a *numerical score*. The purpose is to *rank the fliers correctly*. Numerical score is properly used only for this purpose. The absolute value of the score is (or at least should be) irrelevant.

Besides, in the typical example, what statistical measures will do is to indentify deviations from the norm - so one "good judge" in with several "bad judges" will get eliminated. Never mind that you have still steadfastly ignored the fact that you assume that there *are* errant scores. This is close to oxymoronic, given that you get multiple judges in order to *solicit their opinions*. While you propose ignoring some of them based on mathematical BS that doesn't appear to apply and you have not bothered to prove apply.

If you don't post proofs, derivations, logical arguments, and your premises, you are pretty much admitting you don't know what you are talking about. But of course, you can't, since (and this is the part you should try and understand, if you are capable):

THERE IS NO OBJECTIVELY "RIGHT" ANSWER - AND IF THE SCORE IS JUST A MATTER OF OPINION, YOU CAN'T PROVE THOSE OPINIONS TO BE "WRONG" IN ANY MATHEMATICAL WAY.

There's no rational counter-argument to this point - it's unassailable. But I'm sure that you will continue anyway.

Brett

godzilla · Oct 16, 2004 01:18 AM

#24 source
>The judges are supposed to be using the same criteria for
>judging. If they are, their scores from maneuver to
>maneuver should at least go in the same direction- up is
>better, down is worse. If one judge's scores do not "follow
>the crowd" and track with the other judges, that judge
>cannot be using the same scoring criteria. If he views the
>maneuvers as all "of similar quality" and three other judges
>thought the squares were noticeably better than the round
>loops there is a big problem with the judge's training.

EVERYONE READING THIS THREAD!!!!!

Read every response by Ferocious. He has the idea of the experiment.

This is a measure of the *objective* not the *subjective*.

Hypothesis:
1. There is a perfect manuever. A 40 point score on every manuever is possible. I have a simulation program that flies a perfect pattern. Every manuever is *exactly as described in the rulebook* with ZERO deviations. There are ZERO mistakes and the geometry is PERFECT. A competent judge would be forced to give the simulation a 40 on every manuever.
2. Every judge is grading against this perfect manuever.
3. Every judge is assigning some kind of downgrade for each deviation from the perfect manuever. It really does not matter the scale.
4. In a group of all exceptional judges, the scores would COORELATE. They would not be the same, but there would be a DEFINABLE relationship between the scores (manuever to manuever). Considering that the variation from the ideal 40 point manuever would be constant, even though though the maginitude would be more individual.
5. There are "exceptional" judges and there are "less exceptional" judges. Be this talent or whatever. There certainly would be no argument at a contest that Paul Walker flew *closer to the rulebook* than say JOE ADVANCED FLYER, and his scores would reflect this fact. Judges are certainly no different. Some judges would score more closely to *ideal* scoring, and some would not. It is that simple.
6. Simple 1st order relations appear to prove nothing (by my tests). This would correspond to the graphs of simple linear realtionships as displayed on a simple line graph--ala Gary McClellan. There may be trends, but nothing conclusive (that I can find from the puzzle).
7. The assumtion is that to be conclusive p must be >.05 (probabilty of attaining strange data even if hypothesis is true--- so 1 in 20, this is a standard assumption).
8. The model data set of 600 data points is significant. (I do not have the "rule of thumb" in my notes, but I know that 600 data points is significant to determine the relationship between 4 judges. I will call my Master Black Belt to get the data set limitation numbers).

Last assumptions:

There are some smart CATS out there!
This is a good puzzle.
This is no dig on any judge!!!!!!
This is all made up, and meant to be a learning experience.
I am curious. I have some cool new stats software. That's all!!!!!

N42222 · Oct 15, 2004 10:17 AM

#5 source
Brad,after last year's Nat's and at this year's Tom Farmer Stunt Clinic, Gary McClellan took time from his busy schedule and presented to the DMAA just how the judge's are tracked and rated. He presented slides with judges' individual scores imposed over the others. His slides have the names of judges and flyers removed, of course, and we only see about five or so, but they do this for all the judges and all the flyers. Just another part of the tabulation job, I guess. So everyone at these meetings is "privy". Hope you get an "A" in your class. BOM

ama21835 · Oct 15, 2004 10:39 AM

#6 source
Paul Smith

Back in the late '70's somebody draw a graph of the F2b judges scores vs. the 15 flyers from the home lands of the judges.

As I recall there were US, USSR, French, Italian, and English judges.

The graph looked pretty smooth, except for Pinochio's nose ( the results of French judging French).

ferocious · Oct 15, 2004 12:37 PM

#9 source
looks like you have one part of the problem nailed Brad. Please send me your spreadsheet.

"Also, I am interested to know what methods have been used, or are being used to rate the performance of judges ...."

the only thing I've seen recently was for the World Champs.
http://www.clstunt.com/htdocs/dcforum/DCForumID1/9967.html is a thread that goes into a lot of conflicting detail about judging and scoring at the WC's.

The FAI comparisons involved were at:
http://www.slovanet.sk/orsia/Sheets-F2B_WC2004.xls

and

http://www.slovanet.sk/orsia/WC2004_F2BAnalysis.doc

These are just a very cursory, "look at the scores" discussion which may have upset some people involved. The FAI K factors and the practice of throwing out the high and low scores greatly complicates what is going on in the judging. There is no detail at all about how scores varied by maneuver and judge, so no real comparisons are possible. The original score sheets are probably thrown away or unavailable by now.

Brett Buck · Oct 15, 2004 12:50 PM

#11 source
>looks like you have one part of the problem nailed Brad.
>Please send me your spreadsheet.
>
>"Also, I am interested to know what methods have been used,
>or are being used to rate the performance of judges ...."
>
>the only thing I've seen recently was for the World Champs.
>http://www.clstunt.com/htdocs/dcforum/DCForumID1/9967.html
>is a thread that goes into a lot of conflicting detail about
>judging and scoring at the WC's.
>discussion which may have upset some people involved. The
>FAI K factors and the practice of throwing out the high and
>low scores greatly complicates what is going on in the
>judging. There is no detail at all about how scores varied
>by maneuver and judge, so no real comparisons are possible.

Yes. Even though I don't think either you or Brad has addressed any of the fundamental issues that make the statistical approach invalid, FOR SURE the Bruno Delor thing is almost complete nonsense in terms of figuring anything out about the judging. If it suggests anything, it suggests that tossing scores based on the raw value is an incorrect approach. The basis of this is that there is some "right" absolute value of the score - and that's just incorrect. But of course the whole rest of the analysis is based on this as well.

I repeat - I don't think there is ANY valid statistical basis for deciding which scores to keep and which (if any) to toss, because they aren't proven to be random, and even if they were, the sample size is far to small for most normal statistical measures to be valid. Just because you can put numbers in the equations and calculate the results doesn't mean they are valid.

Brett

klelmore · Oct 15, 2004 12:59 PM

#12 source
If what you're looking for is a way to throw out a judge's score in real time, based on some statistical method, you're right. It is, however, quite possible to examine the posterior scores and determine if one judge is scoring very differently from the others. This assumes, of course, that the scores of the other judges are statistically similar. At that juncture, a CD would have to decide what to do with the end results (adjust all scores by removing the errant judge's scores?) and that is waaaay outta my bailywick.

That errant judge could, however, be tagged for additional training and prevented from judging until that training is completed.

Kim Elmore

Brett Buck · Oct 15, 2004 01:53 PM

#14 source
>If what you're looking for is a way to throw out a judge's
>score in real time, based on some statistical method, you're
>right. It is, however, quite possible to examine the
>posterior scores and determine if one judge is scoring very
>differently from the others. This assumes, of course, that
>the scores of the other judges are statistically similar. At
>that juncture, a CD would have to decide what to do with the
>end results (adjust all scores by removing the errant
>judge's scores?) and that is waaaay outta my bailywick.
>
>That errant judge could, however, be tagged for additional
>training and prevented from judging until that training is
>completed.
>

Of course it's quite easy to pick out scores that don't "track" ex post facto. Gary's method of plotting them and seeing how they go up and down is one way. A perhaps more solid mwathematical way is to normalize the scores from each judge for the whole round, the calculate the deviation at each point. But the assumption that this mistracking is "errant" is not a proven proposition. In fact it's completely indistiguishable from one judge simply preferring the flight for perfectly legitimate reasons.

And for sure using assumptions based on the scores being the result of random variation from some "correct" value is not fundamentally valid. For reasons discussed at extraordinary length before.

Brett

Dr Spark · Oct 15, 2004 02:12 PM

#16 source
I am considering legally changing my name, getting about $50,000 in plastic surgery, ordering new clothes from Harrods of London, and then entering another contest. This to see if my scores will improve without the "baggage" of my former flying.. which seems to follow me.

Dr. Spark

ty marcucci · Oct 17, 2004 12:40 PM

#40 source
Ha, ha, ha. I love this idea. Bs sure to change planes too.

david eyskens · Oct 15, 2004 02:51 PM

#17 source
Just a thought:
Attempting to "quantify" a process, that includes "qualtative" variables, that are "impossible" to "control"-(I think Brett you stated that) seems problematic from the start. Introducing scientific principles into the judging process, although very interesting-(I say this with complete understanding and valadation regarding the OG thread)-may not be the appropriate direction to undertake..The subjective nature of judging is present, and always will be-(maybe I am wrong, but it seems that the goal/purpose of introducing scientific principles into judging would be to reduce the subjective variables related to the process?)..The event has come a long way in improving how to objectify a very difficult resposniblity called judging...The thoughts and ideas arond this issue are interesting, and I respect the individuality concerning it...Thanks David

50plusAirYears · Oct 15, 2004 06:29 PM

#18 source
Bad enough dealing with my 6 Sigma Green Belt assignments at work, now I got to read about Black Belt projects here?
Out of curiosity, do the deviations stay with a given judge throughout the course of the day, or do the judges results show deviation throughout the course of the day with judges exchanging positions with respect to their apparent accuracy?
Actually, I find it interesting to provide assitance on some of our Engineering department Black Belt projects, but spare me from dealing with similar projects from finance or inhuman resources, or other groups that are less "Real World" focussed and concentrate on the "Esoteric Concepts".
We really have made impprovements in our projects processes and Product reliability from the 6 Sigma bit, but we've also had some dead-end projects to. The system definitely has a place in the overall toolbox.

klelmore · Oct 15, 2004 10:38 PM

#19 source
Please understand that my interest at this point is professional. I have no clear notion of why this subject seems to spark even this much interest, except that at some time in the past something notably outrageous happened at a major contest...

> Of course it's quite easy to pick out scores that don't
>"track" ex post facto. Gary's method of plotting them
>and seeing how they go up and down is one way. A perhaps
>more solid mwathematical way is to normalize the scores from
>each judge for the whole round, the calculate the deviation
>at each point.

This depends a bit on the "normalization" used; I'm assuming you mean Z-scores. I'm not sure we'd wantto normalize the scores before we take the residuals (defined here as differences between judges), but we might want to normalize the residuals themselves. I would certainly start with residuals straight away. Raw scores don't possess amenable distributions because, among other things, they're one-sided. Residulas (differences) benefit from being two-sided. In such a case, we have three sets of differences, and in this scenario I suppose we'd look at the correlation structure of the differences. We'd certainly need to do a careful EDA to make understand the residual disributions.

> But the assumption that this mistracking is
>"errant" is not a proven proposition.

I stand (sit? type?) corrected! "Errant" is a poor choice; a better one is simply "different." In this case, whatever statistical model desribes the "different" judge will be different from the model that describes the other judges. Note that we don't need to know what these models are, only that they're different.

>In fact it's
>completely indistiguishable from one judge simply preferring
>the flight for perfectly legitimate reasons.

This may *not* be true in general. I think it's unlikey that one judge's esthetic would be at odds with the others for all fligts, though I suppose it's possible. With enough data (always the gotcha),I'll bet that you can discover that one judge really wants hard, sharp, energy-eating corners while another prefers more graceful corners. We'd have to look at the stats for each individual maneuver (whihc we'd treat as separate factors) to see this, but I'm sure we could do it.

>
> And for sure using assumptions based on the scores being
>the result of random variation from some "correct" value is
>not fundamentally valid. For reasons discussed at
>extraordinary length before.

I'd like to see the discussion -- not to challenge it, but to understand the problems involved *here,* in this application. Would you direct me to the appropropriate thread?

To understand the repeatability of a judge's scores (which is the random variation part), we'd need them to judge exactly the same flight several times under identical conditions, and that's not a possible experiment. A fundamental problem buried in this whole discussion is that we cannot know the source of the variance we observe, much like the problems faced by the social sciences. In the physical sciences (my area), our interest often lies in identifying the variance source.

Brett, are you a statistician, or someone who must use statistics a lot in their work?

Kim Elmore

Brett Buck · Oct 18, 2004 04:59 PM

#76 source
>Please understand that my interest at this point is
>professional. I have no clear notion of why this subject
>seems to spark even this much interest, except that at some
>time in the past something notably outrageous happened at a
>major contest...
>
>> Of course it's quite easy to pick out scores that don't
>>"track" ex post facto. Gary's method of plotting them
>>and seeing how they go up and down is one way. A perhaps
>>more solid mwathematical way is to normalize the scores from
>>each judge for the whole round, the calculate the deviation
>>at each point.
>
>This depends a bit on the "normalization" used; I'm assuming
>you mean Z-scores. I'm not sure we'd wantto normalize the
>scores before we take the residuals (defined here as
>differences between judges), but we might want to normalize
>the residuals themselves. I would certainly start with
>residuals straight away. Raw scores don't possess amenable
>distributions because, among other things, they're
>one-sided. Residulas (differences) benefit from being
>two-sided. In such a case, we have three sets of
>differences, and in this scenario I suppose we'd look at the
>correlation structure of the differences. We'd certainly
>need to do a careful EDA to make understand the residual
>disributions.
>
>> But the assumption that this mistracking is
>>"errant" is not a proven proposition.
>
>I stand (sit? type?) corrected! "Errant" is a poor choice;
>a better one is simply "different." In this case, whatever
>statistical model desribes the "different" judge will be
>different from the model that describes the other judges.
>Note that we don't need to know what these models are, only
>that they're different.
>
>>In fact it's
>>completely indistiguishable from one judge simply preferring
>>the flight for perfectly legitimate reasons.
>
>This may *not* be true in general. I think it's unlikey
>that one judge's esthetic would be at odds with the others
>for all fligts, though I suppose it's possible. With enough
>data (always the gotcha),I'll bet that you can discover that
>one judge really wants hard, sharp, energy-eating corners
>while another prefers more graceful corners. We'd have to
>look at the stats for each individual maneuver (whihc we'd
>treat as separate factors) to see this, but I'm sure we
>could do it.
>
>>
>> And for sure using assumptions based on the scores being
>>the result of random variation from some "correct" value is
>>not fundamentally valid. For reasons discussed at
>>extraordinary length before.
>
>I'd like to see the discussion -- not to challenge it, but
>to understand the problems involved *here,* in this
>application. Would you direct me to the appropropriate
>thread?

http://www.clstunt.com/cgi-bin/dcforum/dcboard.cgi?az=show_thread&om=10256&forum=DCForumID1&viewmode=all

Note that this starts out with me somewhat "primed" as we had already had the identical discussion at least twice before - to the same effect.

>To understand the repeatability of a judge's scores (which
>is the random variation part), we'd need them to judge
>exactly the same flight several times under identical
>conditions, and that's not a possible experiment. A
>fundamental problem buried in this whole discussion is that
>we cannot know the source of the variance we observe, much
>like the problems faced by the social sciences.

Bingo once again.

>Brett, are you a statistician, or someone who must use
>statistics a lot in their work?

Not a statician - but I do deal with a lot of estimation theory in my work (spacecraft attitude determination and control system design). For example, we assess the effects of various noise sources in the sensors to determine the likely worst-case attitude determination accuracy. I just got done checking and approving an analysis of a Kalman Filter system that determines our attitude from gyros, and Earth sensor, and a sun sensor.

Brett

klelmore · Oct 19, 2004 10:34 PM

edited#86 source
Kewl! We use Kalman ensemble filters for predictability studies and ensemble model prediction. My interets lie not so much in generating and running the ensembles, but in evaluating the results.

I worked an example for Bradley that was completely non-parametric using permutation tests. It only showed things about the aggregate scores for each judge, based on the difference in variance between each pair of judges (1v2, 1v3, 1v4, 2v3, 2v4, 3v4). I also showed thet three of the judges generated scores with similar distributions, while the "different" judge generated scores with a bimodal distribution. All I showed was that three of the judges were generating scores that had the (statistically) same variance, while the other judge was using a scoring process that generated much less variance. I also looked at cluster analysis (showed something similar, but lacks a graceful significance level estimate), and PCA using various similarity metrics.

However, there is no way to evaluate whether or not any judge os right or wrong, only that they use a consistent method of generating variance within their aggregate scores (or not).

As you stated in the previous thread, this entire discussion boils down to an exercise in *ranking,* rather than the numerical value of raw scores. The good news here is that rank is a robust statistic that is nearly distribution-free.

Kim Elmore

godzilla · Oct 16, 2004 01:20 AM

#25 source
> And for sure using assumptions based on the scores being
>the result of random variation from some "correct" value is
>not fundamentally valid. For reasons discussed at
>extraordinary length before.
>
> Brett

40 points for perfect with zero flaws is correct. Always.

That is the starting point.

godzilla · Oct 16, 2004 01:44 PM

edited#27 source
> I repeat - I don't think there is ANY valid statistical
>basis for deciding which scores to keep and which (if any)
>to toss, because they aren't proven to be random, and even
>if they were, the sample size is far to small for most
>normal statistical measures to be valid. Just because you
>can put numbers in the equations and calculate the results
>doesn't mean they are valid.
>
> Brett

On what exactly are you basing these statements? Have you done actual studies on scores, or is this simply your standing opinion. If you did study scores, what did you find?

Statistics cannot apply to judging? That is not what I have found, or my Master Black Belt found either. There are several ways to study the data, and in several different directions.

ferocious · Oct 16, 2004 09:05 PM

#29 source
Brett, I believe you are fighting against pretty overwhelming evidence that the task we ask stunt judges to do(estimate a response from a sensory input) does result in random variation. Human judges are not perfect at this task. If you want to throw out statistics as a method of studying scoring and judging I think you need to show some sort of evidence that (by some miracle) stunt judges don't have random variation in their scores.

I don't think Brad's excercise is aimed at browbeating judges or trying to pick some of the scores to throw out(almost always a bad idea). The point is to study the scoring process and figure out how to do it better. Once you identify judges who have a hard time scoring properly then you know who needs more training. Maybe that judge is just using poor techniques like watching the whole maneuver and then assigning a score instead of mentally keeping track of how the maneuver is going as it progresses.

godzilla · Oct 16, 2004 09:28 PM

#31 source
>I don't think Brad's excercise is aimed at browbeating
>judges or trying to pick some of the scores to throw
>out(almost always a bad idea). The point is to study the
>scoring process and figure out how to do it better. Once
>you identify judges who have a hard time scoring properly
>then you know who needs more training. Maybe that judge is
>just using poor techniques like watching the whole maneuver
>and then assigning a score instead of mentally keeping track
>of how the maneuver is going as it progresses.

On the nosey...

It would be a useful tool to determine who is having trouble.

Brett Buck · Oct 17, 2004 10:54 PM

#49 source
>Brett, I believe you are fighting against pretty
>overwhelming evidence that the task we ask stunt judges to
>do(estimate a response from a sensory input) does result in
>random variation. Human judges are not perfect at this
>task. If you want to throw out statistics as a method of
>studying scoring and judging I think you need to show some
>sort of evidence that (by some miracle) stunt judges don't
>have random variation in their scores.

No, it's your concept, you have to show it, and what distribution it follows. And what parameter it might be that contains the random element, And develop statistical measures that are appropriate to the distribution, and for exceptionally small sample populations. You have done none of this, nor have you even begun to prove mathematically what you mean. Please go ahead - I can probably understand it if you explain it properly. I suggest you start with answering some of the very fundamental questions raised previously - instead of studiously ignoring them, waiting a while. and then repeating the same assertions.

Or alternately, you can just state and restate the same point over and over without proof.


Brett

ferocious · Oct 18, 2004 05:00 PM

#77 source
according to all the articles and texts I've read, the judge's scoring variations would be approximately on a Gaussian distribution. You would probably see some peaking at the 5 point intervals and more at the 10 pt intervals(people tend to like certain numbers). You'll also see some avoidance of the high and low score. People don't like to use the extremes. But overall, that doesn't affect the analysis enough to bother with.

As Kim points out, you'd want at least 10 judges for reasonably good scores, preferably 25-30, to get the estimated error down below 1 pt.

It doesn't matter that the score is an "opinion". If an opinion is what you are looking for, use only one judge. Then there can be no questions about scoring because the one judge's opinion is the only one that counts. A multitude of opinions become a fact, as in opinion polls, taste panels, and elections.

I heartily applaud Brad for putting his puzzle out there. Brave man.

ty marcucci · Oct 15, 2004 11:07 PM

#20 source
Very interesting. Benjamin Disralei, the British Prime Minister of long ago once said, "There are lies, ##### lies and statistics". The one factor left out, or so it seems, is the human element, which seldom if ever is measured. How tired was judge "X", what is the vision of judge "Y", what happened to judge "M" last night, what is wrong with judge "T"s glasses, Why does judge "Q" hate green planes, why does judge "L" dislike the flyer so much?? ad infinitum.
All this being asked, I hope your system comes up with something useful for CLPA, and it seems you sure do come up with good stuff and think out of the box. Keep it up.

DMoon · Oct 16, 2004 12:00 AM

Cant use stats on judegs!edited#21 source
Using stats to determine a judge's competence is not really valid in any sense. Simply because the judges dont weigh the portions of a maneuver the same. If they did move up and down the same amounts then we would only need one judge at the contests right??

I used to think stats would be a viable way but after the following example I no longer think it would be worth the time involved in tracking. The only REAL thing it would prove is that they dont all score the same. We already know that.

In the last contest I flew overhead eights on my first flight.

Personally I thought they werent so great. I thought about a 30 would be a good score.

I get my score card and one judge has me at 28 and the other has me at 34! These two judges are very experienced. They both have seen me fly many times. One has judged at the Nats OTS and Classic and INT many times and both fly at the Nats every year. These two have lots and lots of experience looking at and scoring patterns.

BUT why the large difference in the score?

Well the low judge looked at it like I did. Here is how he and I saw it.

Missed the int twice across the top. Inside loop not straight above my head with the outside loop.

(He and I and many others tend to focus on the mistakes only!)

Here is how the other judge saw it.

Entered it striaght above my head. Exited it straight above my head, same as entry. He said I was one of the few who actually flew the overheads at the 45 mark, most others went well below. He did see my other mistakes but he also saw how I was trying to fly the maneuver within the correct parameters making it more correct, according to him.

They both saw the same maneuver. They both saw the same mistakes, however only one saw the GOOD things in the maneuver.

The stats would say the high judge was off his rocker but he did have his reasons, good ones I might add, for scoring it high. Not just some guy writing numbers down.

This cant be accounted for with the stats stuff.

And that is why I think Stats is an invalid way to track judges. Comparing them to each other without them being able to explain why they scored the way the did doesnt give them a fair shake.

Peabody said he was going to start judging is Hourglass differently. He says what he is seeing these days isnt rulebook. Well the stats will say he is a kook.

He was saying that the guys dont fly the top and the bottom the same length. I broke out the videos of the WC and the Nats04 and the Nats 02. Got out my dry erase marker and started in on the screen. First off there were very few that were at the very perfect spot in front of the camera for me to mark off on the screen.(The judges were NOT ususally in the right spot either, pretty close but that can scue THAT maneuver alot) But of the ones the camera did get, the fliers got the bottoms and the tops pretty DARN close to the exact same length on my TV screen. I even measured one guy and it was dead on. They seem to fly it pretty darn close to correct if you ask me.

On another note I dont think bracket scoring is as wide spread as some think. We practice our patterns as a whole. They will progress as a whole. The inside and outside rounds often look very close to each other. Same with the squares. That is 4 maneuvers that will score very similar to each other on the sheet. Also many people think they take off, level flight as well as anyone. WORNG! I watched a whole group of expert fliers at the labor day contest, 8 or 10 fliers and only ONE flew the proper take off, ONLY ONE! That is why the scores stay in the 30 range. AND I over heard one judge say that very thing! The fliers walk around wondering why but it is on them. I wonder myself alot. But in the end it is on me.

Also, about inverted flight. It does not say anything about the plane not deviating from a perfect straight line. Yet many judges use this as a criteria to get a perfect score. If you read the book it says you have to stay in the 4'-6' range and that should be perfect inverted flight. Seems like you can wonder up and down a bit but you cant get out of the 4'-6' range, according to the book. I remember Johnny D brought this up once at the Nats pilot's meeting and people looked at him like he was crazy. I think it is a good point. What gives?

Sorry to ramble on so. I havent been on here in a while.

godzilla · Oct 16, 2004 12:33 AM

RE: Cant use stats on judegs!#22 source
>Using stats to determine a judge's competence is not really
>valid in any sense.

Careful now. All of the things mentioned in your reply are taken into account of a good analysis.

Statistics are not simple. they really are not. They are really quite mind blowing to me.

godzilla · Oct 16, 2004 01:28 AM

RE: Cant use stats on judegs!#26 source
>Peabody said he was going to start judging is Hourglass
>differently. He says what he is seeing these days isnt
>rulebook. Well the stats will say he is a kook.

He would be an outlier on ONE manuever. Over a data set of even a few fliers, this would be insignificant.

This is to Brett's point, that one bad manuever score would be nearly impossible to find (needle in a hay stack---even though... if the data could be compared to a KNOWN GOOD it would not be hard).

I am not interested in that. I am proposing the WORST CASE SCENARIO. Either a judge that is "bracketing" or simply scoring WRONG for the entire pattern or "faking it".

klelmore · Oct 17, 2004 10:38 AM

RE: Cant use stats on judegs!#37 source
Everything you discuss in your post positively *CAN* be evaluated statsitically, but not at any single contest. Judges' comments can be categorized by whatever characteristics interests the analyst, and the the *comments* can absolutely be evaluated statistically. This sort of thing is done all the time. It's called a factor-based analysis -- which shouldn't be comefused with Factor Analysis, a category of a general family of eigen techniques first developed by Hotelling in the '30's.

Kim Elmore

godzilla · Oct 16, 2004 02:03 PM

#28 source
Here is the spreadsheet with the data.

If you solve it, you must show your work!!!

DMoon · Oct 16, 2004 09:33 PM

#32 source
I dont know about all of this stuff. Very interestig for sure. But the real crux of the problem is the way the rulebook states what is what. That meaning there is not one part of each maneuver that is worth more than the other. Size, Shape, Corner and so on. The way I have read it they are all to weigh the same. Yet I hear one judge say if it is too big there simply is no shape. POW you go big in front of that guy and you are in the tanker. Another says if it is a little big but all the sides are good and the bottoms are proper height then big is the only ding. Another guy wants to see good bottoms above all else simply because it is easy to see. You see where I am going with this.

Your stats while fun and VERY Interesting would really show nothing without explaination from the judges as too why they score the way they do.(Dont forget there is a huge Human Element here) You pull a judge because his stats say he is a coaster when in fact he uses a criteria, well within the rulebook, that is different than the norm. Come up with the "NORM" first and then your stats will prove out every time. Now it is just tagging judges who judge different, but not worng, than the rest of the panel. See my example in my previous post.

Only when the criteria is really set in stone on what is worth the most and down from there will the judges be able to tracked by the stats with any kind of validity AND the fliers will actually know what they are in for no matter where or who they fly in front of. IF it were really defined in a straight forward manner then you would see the judges move in the same direction. You would also see a surge in what is really scored well being practiced more. The way it is now it is a crap shoot. So you better practice everything!

50plusAirYears · Oct 16, 2004 10:39 PM

#34 source
Unless there is some kind of measurement system like the stopwatch in FF or in speed and carrier, judging is principally subjective, therefore will automatically have a lot 'noise' in the data.
Also, the concept of a judge giving favorable scores to a countryman is not restricted to our sport. It has been a problem with other sports, such as in the Olympics for decades.

Pat Mackenzie · Oct 16, 2004 11:41 PM

#35 source
Based on your description of the problem and my analysis of the first round data, Judge 1&2 are seeing the same thing. , judge 4 is "braketing" (If I understand the term) and judge #3 is not seeing the same thing as 1&2.
Pat MacKenzie

Pat Mackenzie · Oct 17, 2004 11:21 AM

#38 source
I didn't use any complex statistical methods, just followed Ross Perot's lead and stuck with "Charts and Graphs"
Here is one graph. One flyer for one flight, all four scores. Scores for 1&2 track fairly well, which was your definition of "correct judging" for this experiment. Judge 3 doesn't track the other two very well and judge 4 just wanders up and down.
Pat MacKenzie

ty marcucci · Oct 17, 2004 12:20 PM

#39 source
Judge four reminds me of a judge that stated out loud that "a person flying in Intermediate can't get over 25 to 30 points per maneuver". This really baffles me. If so, then how does one get into Advanced, if this judge is the one always judging Intermediate?

Never fly at just one particular contest, move around. I have had flights as low as 290 and as high as 425 in Intermediate, but at extremes in distance from each contest. Same plane, same me, same engine, same fuel, same prop, not all that much more experience and the same summer season. These were complete flights, Crashes due to stupidity, bad fuel, poor flying, etc don't count here.

During the Navy sponsored NATS, the judges were all Naval Aviators, officers, with little or no experience in judging until a few days prior to the event. So training them is not all that hard.

More stunt judging clinics are definitely needed. I have heard of only two in the last five years. Comments?

godzilla · Oct 17, 2004 01:42 PM

#41 source
Kim,

The trends do show some coorelation, but they do not prove anything. I think a graph like this one is inconclusive even though it could be said to be somewhat obvious. I am looking for more bulletproof methods that could be used if the data was even less obvious.

It does go to show however that the graph like you used could be used to make an argument one way or another if the data were obvious.

It does not appear that you used all of the availble data. It appears you only used one flight score to solve the puzzle. This would be an invatation a for fist fight, as no one would evaluate anyone's performance from one flight score. I know I would not think this proves anything.

Solving the puzzle is not as important as finding a suitable method that can be used over a large data set. Maybe I should have been more clear. The method is the thing I am interested in, the solution is less important.

Thanks for playing though! See if you can plot the entire data set and see the same coorelation, I doubt you will be able to.

klelmore · Oct 17, 2004 02:46 PM

edited#44 source
Wasn't me that posted this example, it was Pat. I commented, however.

Pat Mackenzie · Oct 17, 2004 04:59 PM

edited#46 source
>The trends do show some coorelation, but they do not prove
>anything. I think a graph like this one is inconclusive
>even though it could be said to be somewhat obvious. I am
>looking for more bulletproof methods that could be used if
>the data was even less obvious.

>It does not appear that you used all of the availble data.
>It appears you only used one flight score to solve the
>puzzle. This would be an invatation a for fist fight, as no
>one would evaluate anyone's performance from one flight
>score. I know I would not think this proves anything.
>
>Solving the puzzle is not as important as finding a suitable
>method that can be used over a large data set. Maybe I
>should have been more clear. The method is the thing I am
>interested in, the solution is less important.
>
>Thanks for playing though! See if you can plot the entire
>data set and see the same coorelation, I doubt you will be
>able to.
Actually I graphed the entire first round, and the trend was very clear. There was no need to plot the second round. You had already stated what "errors" you had hidden in the data, and they really do jump out. The biggest thing was how well the two "good" judges track, which is the underlying assumption you started with.
While your "six sigma black belt" guru might not like the non-statistical nature of my solution, I would argue that it sometimes takes a lot of analysis to quantify something for which the trends are clear if data is just plotted.
And of course there are two old expressions that apply to statistics:
"Figures lie, and liars figure" Unknown author.
and of course
"lies, darned** lies, and statistics" Mark Twain
Pat MacKenzie
** not the right word, but Len's software filtered the other one

godzilla · Oct 17, 2004 05:35 PM

#47 source

>Actually I graphed the entire first round, and the trend was
>very clear. There was no need to plot the second round. You
>had already stated what "errors" you had hidden in the data,
>and they really do jump out. The biggest thing was how well
>the two "good" judges track, which is the underlying
>assumption you started with.

Once agin, this graph would prove nothing in a Real World setting. I would be very surprised if anyone would notice the trend if they had not been SPECIFICALLY LOOKING for it.

A good test would work for all data sets.

Not to say you did not find the solution, but you would have a hard time convincing a skeptic.

klelmore · Oct 17, 2004 02:43 PM

#43 source
A good graph that makes an interesting point. Here's a thought, and perhaps a way to think *statistically* about what we're at here. NOTHING presented by me or anyone else defines *correct* judging. In When the trus values cannot be known, all statsitics can tell us about is whether or not judges are consistent. We cannot address correctness or its lack. We can only address consistency and similarity. If we address correctness, then we enter the realm of validation, and we have never been able to do that with nothing more than subjective scores.

In your example, there is no way to show whether judges 1 and 2 are correct compared to judge 3. We can only show that 1 and 2 are scoring consistently compared to each other, while judge 3 is scoring based on different criteria.

Kim Elmore

godzilla · Oct 17, 2004 02:55 PM

#45 source
>A good graph that makes an interesting point. Here's a
>thought, and perhaps a way to think *statistically* about
>what we're at here. NOTHING presented by me or anyone else
>defines *correct* judging. In When the trus values cannot
>be known, all statsitics can tell us about is whether or not
>judges are consistent. We cannot address correctness or its
>lack. We can only address consistency and similarity. If we
>address correctness, then we enter the realm of validation,
>and we have never been able to do that with nothing more
>than subjective scores.
>
>In your example, there is no way to show whether judges 1
>and 2 are correct compared to judge 3. We can only show that
>1 and 2 are scoring consistently compared to each other,
>while judge 3 is scoring based on different criteria.
>
>Kim Elmore

That is true. This could be very useful however in addressing each individual judge's training. If this result were over say an ENTIRE Nats I think the data would be tremendously useful. In fact, it would probably be just as easy to say that a null hypothesis of zero coorelation between judges would be as valid as saying the judges SHOULD coorelate (especially if you are in the "majority subjectivity" group).

Since I am not one of those people (I believe judging should be a "majority objective" endeavor---I explain in the other post). I believe that if the scores do not coincide beween judges, we HAVE A PROCESS OUT OF CONTROL.

***A process out of control would have low predictability, possibly ZERO predictability.***

I know of not one flier who would except the fact that the outcome of a contest should be "subjectively random". I can honestly say that I do not know one single flier who would say this. Just the opposite is true, every flier I have ever spoken to feels that the BEST flier should win, most of those would say the BEST FLIER AS DEFINED BY THE RULEBOOK should win.

I believe in the 80-20 rule.

godzilla · Oct 17, 2004 02:42 PM

Scoring Methodedited#42 source
>Your stats while fun and VERY Interesting would really show
>nothing without explaination from the judges as too why they
>score the way they do.(Dont forget there is a huge Human
>Element here)

The assumption would be that all of the judges had recieved the same training and are basing their scores on the same range of deductions. Today, I am not sure you can say this is true. I say I am not sure, because I am not sure how the judges are trained. I would hope that the judges are trained to look for the most common mistakes and the suggestion from the trainer would be that the deductions should coorelate.

My assumptions are that there is a percentage of "objective" and "subjective" criteria, each individual to each judges perspective. I would like to see a ratio of about 80% objective and 20% subjective criteria. I have heard arguments that our event is 100% subjective, and 100% objective. Depends on what side of the fence the person is sitting on, on that particular day. Some argue that the rulebook is a "guideline" and that the manuever score should be left open to some kind of interpretation from each judge. I have also heard it loudly argued that the rulebook is the MEASURE of each manuever it is an absolute measure. I think both approaches, if taken on the whole, are incomplete. Like running all castor or all synthetic oil, sometimes the best idea is a BLEND.

Let me explain how I judge as an example. I have a very simple system that works for me, and I have a few friends here in Texas that have shared ideas on this method and use it with only minor twists in the deductions from judge to judge.

1. The number one objective criteria is shape. If the shape is wrong, for several reasons, the deduction counts double over a simple bobble or wiggle. Shape errors are obvious to anyone would looks at the rulebook description so I will not into the detail. A shape error counts two points per occurrance except in the instance of the wingover and the hourglass where they count 4 points per occurrance. This is due to being single manuevers. The clover while being a single manuever has 4 loops and two intersections so a 2 point deduction is sufficient. Each portion of the manuever has its own individual shape error deduction (for example, the square eight has 6 potential incorrect shapes, 4 squares and two intersections). Shape errors include any crooked legs, hooks in rounds or flats, size, and gross intersection mistakes.

2. A bobble is one point for each occurrance.

3. The number one subjective criteria is corner size for angled manuevers. There are 3 catagories that I use. Hard (close to rulebook), medium (most everyone), and soft. The flier would recieve an overall dedcution depending on what catagory he falls under. Hard corners receive no deduction, medium two points, and soft 4 points. These are assessed at the completion of the manuever. The total score would be assessed starting at the new potential score. For the starting point would be 40 for hard corners, 38 for medium corners and 36 for soft corners.

4. The last deduction catagory for me is bottom height. This is such a common mistake that I have made it a standard deduction as it is easier to keep up with. This is reserved for the fliers that fly good patterns but the bottoms are obviously not at the 4 to 6 foot range. I use a deduction of 1 point per foot. So a 10 foot bottom manuever would receive an automatic 4 point deduction so the flier would be starting a 36 point perfect score.

Using this system no particular "style" is at a disadvantage. If your style is to try to fly the pattern EXACTLY as stated in the rulebook, go for it, you will be starting at the greatest potential score! You will also be most likely to make errors over another flier who is compromising shape or corners for precision.

It is likely however that a compromise flight could score better if the excecution is superb. Let say you are flying a soft corner pattern, but your excecution of shape and size are extraordinary. With zero bobbles you could achieve a 36 for any manuever with corners. That is a great score, but you better be PERFECT on shape and not have any wiggles or the score could drop from the already comprimizing 4 point shortage.

As stated before, size is considered a shape error. It is also deducted at 2 points FOR EACH OCCURRANCE. So a flier doing a square eight with perfect medium corners, perfect bottoms, intersections, tops and shapes (geometric shape) would receive a 8 point shape deduction and a 2 point corner deduction. A perfect excecution of the manuever under this criteria would receive a 30.

****I am still undecided as to whether a 2 point deduction for shape for size is fair, I guess it depends on the camp of judges you are talking too, a 1 point deduction per occurance might be a better rule. As long as it is consistance I would not care.*****

DMoon · Oct 17, 2004 09:17 PM

RE: Scoring Method#48 source
>The assumption would be that all of the judges had recieved
>the same training and are basing their scores on the same
>range of deductions.

That is an assumption I would never make...

DING DING DING DING DING!!!!!!!!!!

There are many hours spent at the nats to try to get the judges on the same page but a simple scan of your sheets will show that some people just "Favor" certain things more than others.


>
>Let me explain how I judge as an example. I have a very
>simple system that works for me, and I have a few friends
>here in Texas that have shared ideas on this method and use
>it with only minor twists in the deductions from judge to
>judge.
>
>1. The number one objective criteria is shape.

"Without shape there simply is no maneuver"

Tom Farmer.

If the
>shape is wrong, for several reasons, the deduction counts
>double over a simple bobble or wiggle. Shape errors are
>obvious to anyone would looks at the rulebook description so
>I will not into the detail. A shape error counts two points
>per occurrance except in the instance of the wingover and
>the hourglass where they count 4 points per occurrance.
>This is due to being single manuevers.

So there is a slight change in direction at the top of the WO on both tracks that moves him to 32. He also missed int on the repeat side of the arc, at the begining, that is another 4 to move him to 28. Then he pulls out at 8' on the inverted side so he is at 26. ON the upright side he is at 7' so he is at 25. There was bobble on the outside pullout that moves him 24. I have seen this very WO from many top pilots around here and they arent getting 24s. I know you have gotten a 24 a time or two and not been too happy about it. Thing is the WO I described is VERY Common anywhere you go. VERY few have even bottoms, no direction changes at the top and perfect ints., and no bobbles or very soft inverted pullouts. By your standards Bob Gieseke would very easily be the high judge...

The clover while
>being a single manuever has 4 loops and two intersections so
>a 2 point deduction is sufficient. Each portion of the
>manuever has its own individual shape error deduction (for
>example, the square eight has 6 potential incorrect shapes,
>4 squares and two intersections). Shape errors include any
>crooked legs, hooks in rounds or flats, size, and gross
>intersection mistakes.

Dont forget the repeatablity thing...you are supposed to repeat the first maneuver you do the second time by. So if he blows it the first time but repeats it you are not to kill him on that second shape error. Talk to Big Art about that. Your method completely ignores this. How come?

>
>2. A bobble is one point for each occurrance.

Hard corner with a bobble 1 point. Medium with no deviation 2 points...this rewards the flier trying for RB corners????

>
>3. The number one subjective criteria is corner size for
>angled manuevers. There are 3 catagories that I use. Hard
>(close to rulebook), medium (most everyone), and soft. The
>flier would recieve an overall dedcution depending on what
>catagory he falls under. Hard corners receive no deduction,
>medium two points, and soft 4 points. These are assessed at
>the completion of the manuever. The total score would be
>assessed starting at the new potential score. For the
>starting point would be 40 for hard corners, 38 for medium
>corners and 36 for soft corners.
>
>4. The last deduction catagory for me is bottom height.
>This is such a common mistake that I have made it a standard
>deduction as it is easier to keep up with. This is reserved
>for the fliers that fly good patterns but the bottoms are
>obviously not at the 4 to 6 foot range. I use a deduction
>of 1 point per foot. So a 10 foot bottom manuever would
>receive an automatic 4 point deduction so the flier would be
>starting a 36 point perfect score.

How do you see a 10 foot or a 9 foot bottom at 150 feet away??

Are size errors a deduction in shape as well or do you count them seperate? Like size is a deduction and shape is a deduction. Some feel too large automatically kills the shape. Others say it is simply a separate deduction and look at the shape from there. That is why you sometimes see a larger pattern get killer scores.

I have looked this over a few times. Heck I couldnt get a 25 in front of you on any maneuver. Especially my inverted flight...

When was the last time you judged using this method? You will be the low judge for sure. Is two laps long enough to do all the figuring?

I like the fact that you have come up with a method for assessing a score. I would like to hear what others use. I dont have a method. Your data set would kick me out... But then again I dont judge very often. I do feel I have a good idea of what it should look like and when it doesnt I count down from there.

>
>Using this system no particular "style" is at a
>disadvantage. If your style is to try to fly the pattern
>EXACTLY as stated in the rulebook, go for it, you will be
>starting at the greatest potential score! You will also be
>most likely to make errors over another flier who is
>compromising shape or corners for precision.
>
>It is likely however that a compromise flight could score
>better if the excecution is superb.

Uh what? That totaly contradicts the above paragraph. I seem to remember someone really pissed that they tried to for RB sizes and a larger pattern beat them. I cant believe you would now say that would be ok. FACT IS! If you compromise your pattern then your execution can NEVER be superb. It is already COMPROMISED! That sentence really contradicts itself.

Let say you are flying
>a soft corner pattern, but your excecution of shape and size
>are extraordinary. With zero bobbles you could achieve a 36
>for any manuever with corners. That is a great score, but
>you better be PERFECT on shape and not have any wiggles or
>the score could drop from the already comprimizing 4 point
>shortage.

Should only be a 2 point deduction. He repeats the mistake on purpose ust as he is supposed to. Then add 1 back in for that very good thing that he did correct, the repeating thing.

What happend when said flier flies everything really well size shape bottoms. However he is really going for RB corners and getting bobbles here and there. He should score better since he is trying for RB pattern. The compromise should never outscore. Dont you think so? Does this happen in real life? HECK NO!

That is totaly unfair to the guy busting it to get to the RB pattern and trying for it in compitition. This was this very fliers argument. He was striving for the RB and others were backing off of it and getting ahead. The RB is what we all strive for and those backing off should be dinged for it.

>
>As stated before, size is considered a shape error. It is
>also deducted at 2 points FOR EACH OCCURRANCE.

Remember the repeat the maneuver thing....

So a flier
>doing a square eight with perfect medium corners,

How is a medium corner considered perfect?

perfect
>bottoms, intersections, tops and shapes (geometric shape)
>would receive a 8 point shape deduction and a 2 point corner
>deduction. A perfect excecution of the manuever under this
>criteria would receive a 30.

So he gets a 2 point corner deduction and since they are "medium" they also contribute to a shape error? I dont see where you said he had a shape error. Howcome the 8 point deduction?

But what about the fact that he has repeated the same thing over and over? That has some weight to his advantage. The maneuver you have described would probably score ALOT higher than 30 on most circles. If medium corners are the only issue I will take it every time!

>
>****I am still undecided as to whether a 2 point deduction
>for shape for size is fair, I guess it depends on the camp
>of judges you are talking too, a 1 point deduction per
>occurance might be a better rule. As long as it is
>consistance I would not care.*****

Interesting...

I like that you are showing your method. Shows some very forward thinking. All judges have their OWN methods right? So how in the world can we begin to think they should all fall under the same data set and it would somehow be a fair assessment of the judges ability?

Brett Buck · Oct 18, 2004 12:20 AM

RE: Scoring Method#52 source

>I like that you are showing your method. Shows some very
>forward thinking.

But not clear thinking. It's just gibberish, because:

>All judges have their OWN methods right?
>So how in the world can we begin to think they should all
>fall under the same data set and it would somehow be a fair
>assessment of the judges ability?


Bingo, once again.

By far the soundest method is to take *all* the judge's scores (no throwouts), use their scores only to rank the fliers for each judge, and pick the finishing order based on the accumulated rank. Any method that uses the raw score directly, and compares them from judge to judge, is patently incorrect, since the raw score does not mean anything in and of itself. Phil's delusions of objectivity are just that, delusions.

Brett

godzilla · Oct 18, 2004 08:11 AM

RE: Scoring Methodedited#54 source
>
>>I like that you are showing your method. Shows some very
>>forward thinking.
>
> But not clear thinking. It's just gibberish, because:

Brett, I did not think it was possible, but I think you are becoming more and more abrasive.

Did you just say Phil is delusional?

Is this supposed to be discriptive or cute?

N42222 · Oct 18, 2004 09:47 AM

RE: Scoring Method#56 source
Gabby Walker is RIGHT! Not only was it pure frontier gibberish, it displayed a level of knowledge not heard too often around these parts. I just wish the childern could have been here.

godzilla · Oct 18, 2004 10:27 AM

RE: Scoring Method#57 source
>Gabby Walker is RIGHT! Not only was it pure frontier
>gibberish, it displayed a level of knowledge not heard too
>often around these parts. I just wish the childern could
>have been here.

Ruuruh!!!

Thank you Howard Johnson!!!!

Hurumph! Hurumph!

Dick Fowler · Oct 18, 2004 10:42 AM

RE: Scoring Method#58 source
WOW!.... I'm waiting for - "Brad you ignorant slut!"

godzilla · Oct 18, 2004 11:01 AM

RE: Scoring Methodedited#59 source
> By far the soundest method is to take *all* the judge's
>scores (no throwouts), use their scores only to rank the
>fliers for each judge, and pick the finishing order based on
>the accumulated rank. Any method that uses the raw score
>directly, and compares them from judge to judge, is patently
>incorrect, since the raw score does not mean anything in and
>of itself.


First of all, I never said I would use any of the methods discussed in this thread to determine to outcome of a contest. All I have asserted is that the methods discussed could correlate the performance of one judge to another. This may be useful in determining who is in need of further training, or is simply clueless, tired, or is not trying.

In the case of deciding a contest, your idea has merit. I so wish I had an idea that has some merit, so I could be smart like you!

>Phil's delusions of objectivity are just that,
>delusions.
>
> Brett

I might watch out, Phil is a combat flier, he is not a round belly stunt flier with thick glasses. He might just dot your eye.

Brett Buck · Oct 18, 2004 11:29 AM

RE: Scoring Method#60 source

>I might watch out, Phil is a combat flier, he is not a round
>belly stunt flier with thick glasses. He might just dot
>your eye.

Well, I assume he is also an adult.

Brett

godzilla · Oct 18, 2004 01:27 PM

RE: Scoring Method#61 source
>
>>I might watch out, Phil is a combat flier, he is not a round
>>belly stunt flier with thick glasses. He might just dot
>>your eye.
>
> Well, I assume he is also an adult.
>
> Brett

It would be very adult to administer a spanking to a child with a smart mouth.

wmiii · Oct 18, 2004 01:58 PM

RE: Scoring Method#62 source
Your way over the line here, or are you the adult with the smart mouth?

Walter

Brett Buck · Oct 18, 2004 02:10 PM

RE: Scoring Method#63 source
> Your way over the line here, or are you the adult with the
>smart mouth?

Note that confronted with a challenge that should easily be provable if it was right, the riposte is calling me "four-eyes" and "fatty" (although that one is borderline inexplicable, given my 36" waist size) and then suggesting I need to be punched. Classic 8-year-old playground bully stuff. Guess he forgot "chromedome".

I'm done attempting to discuss anything with Brad. I'll respond only to correct the errors, and not even attempt to respond to him directly, as it's clearly pointless.

Brett

dirtydan · Oct 18, 2004 02:33 PM

RE: Scoring Method#64 source
> Note that confronted with a challenge that should easily
>be provable if it was right, the riposte is calling me
>"four-eyes" and "fatty" (although that one is borderline
>inexplicable, given my 36" waist size) and then suggesting I
>need to be punched. Classic 8-year-old playground bully
>stuff. Guess he forgot "chromedome".
>
> I'm done attempting to discuss anything with Brad. I'll
>respond only to correct the errors, and not even attempt to
>respond to him directly, as it's clearly pointless.
>
> Brett

Brett,

And you recently (mildly and with remarkable taste) chastised me for pointing out that Godzilla's Eddy Haskell persona has made a return.

Uh, is "chromedome" a negative comment? I've gotten that for years, have I been mislead in thinking it means, "We love you, especially your lovely head, unsullied as it is by unruly hair."

Dan

wmiii · Oct 18, 2004 02:47 PM

RE: Scoring Method#66 source
Brett you do know that I was talking about Brad, right?

Walter

Brett Buck · Oct 18, 2004 02:50 PM

RE: Scoring Method#67 source
> Brett you do know that I was talking about Brad, right?


Certainly.

Brett

godzilla · Oct 18, 2004 03:31 PM

RE: Scoring Method#68 source

> I'm done attempting to discuss anything with Brad. I'll
>respond only to correct the errors, and not even attempt to
>respond to him directly, as it's clearly pointless.
>
> Brett

Classic psuedo-intellectual strategy. Beat people up and then call them incapable of carrying on reasonable conversations when they finally tire of the constant bashing. Somewhere along the line the 1st Amendment gives them the right.

Read the thread.

Brett is the constant agressor. He has become more abusive in his choice of words in response to people's opinions, especially in last month or so. Any attempt to defend one's self just escalates the entire thing.

Actually, I never said anyone should hit Brett. I am simply amazed that somebody somebody hasn't considering the way he talks to people.

As far as calling him "fat" or "four eyes", give me a break, I outweigh Brett by 50 lbs. Roundbelly describes most stunt people. It does not describe most combat people.

Once again, this is all in Brett's head. Brett gets caught being a jerk once again and deflects it to someone else.

Brett Buck · Oct 18, 2004 03:54 PM

RE: Scoring Method#70 source
>
>> I'm done attempting to discuss anything with Brad. I'll
>>respond only to correct the errors, and not even attempt to
>>respond to him directly, as it's clearly pointless.
>>
>> Brett
>
>Classic psuedo-intellectual strategy. Beat people up and
>then call them incapable of carrying on reasonable
>conversations when they finally tire of the constant
>bashing.

So asking for a proof is "beating them up"? So, I guess we'll see that derivation of statistics for N=3 any time now? And you'll explain how you objectively determine "truth" and accuracy of a subjective opinion?

And, of course, it makes perfect sense in your world that someone would resort to violence in response to a vigorous academic challenge?

As far as the rest of it goes, I'm perfectly comfortable with letting the records speak for themselves as far as who the posers are in this context.

Brett

godzilla · Oct 18, 2004 04:08 PM

RE: Scoring Methodedited#72 source
> So asking for a proof is "beating them up"? So, I guess
>we'll see that derivation of statistics for N=3 any time
>now? And you'll explain how you objectively determine
>"truth" and accuracy of a subjective opinion?

I never answered the question, so what, there are 20 questions asked you in this very thread you have yet to answer. I have a method, I can show it, that is all I have said. If you have a method or a proof that there is NO POSSIBLE METHOD, please post it. It is not my intention to spend hours defending what I am doing simply BECAUSE YOU SAID I SHOULD.

Also, it was my intention to get more methods than my own. I already gave away too much information about the data sets. I was hoping that a few methods could come back to me then I would discuss what was sent back.

Once again, this IS NOT ABOUT BRETT!!!

> And, of course, it makes perfect sense in your world that
>someone would resort to violence in response to a vigorous
>academic challenge?

"Vigorous academic challenge" I have no problem with. Speaking to me or my friends in a constant demeaning tone, using condescending words constantly, I do have a problem with.

No one said they were going to hit you. You sound like a big girl.

> As far as the rest of it goes, I'm perfectly comfortable
>with letting the records speak for themselves as far as who
>the posers are in this context.
>
> Brett

Now I am a stats poser too? Your self image has now reached global proportions.

I start a thread. Show my work, ask for people to look at the work and I am a poser?

godzilla · Oct 18, 2004 03:43 PM

RE: Scoring Method#69 source
>
> I'm done attempting to discuss anything with Brad. I'll
>respond only to correct the errors, and not even attempt to
>respond to him directly, as it's clearly pointless.
>
> Brett

No doubt. Better yet just don't respond as it is obvious that you have a particualr affinity for any opinion I do not have.

Once again you have done nothing on this thread but state that it is pointless for us to discuss anything, since you already know the solution.

So far, the only point you have made is that the entire excercise is pointless.

godzilla · Oct 18, 2004 03:59 PM

RE: Scoring Method#71 source
> Your way over the line here, or are you the adult with the
>smart mouth?
>
>Walter

Show one place in this thread I said anything smart that was not in response to Brett.

What I do not understand is why people continue to defend his obvious lack of respect for anyone's opinion except his own. I could certainly carry on a thread on SSW wih ZERO input from Brett, considering it is all negative.

ferocious · Oct 18, 2004 05:31 PM

RE: Scoring Method#80 source
you will get exactly the same "errors" if you go by rank instead of score. Ranking and Scoring generally give the same results. Most of the articles I've seen show them to be equally variable, when using humans as judges, and there usually isn't any scientific basis to choose one method over the other. However, if the judges are doing a good job, the scores have more info in them than just the ranks. Ranking throws out the absolute differences between the scores which can be useful data. Most people would rightfully think that a 540 pt. flight was quite a bit better than a 330 pt. flight, and a more meaningful difference than a 540 pt flight ranking ahead of a 539 pt. flight.

godzilla · Oct 18, 2004 08:44 AM

RE: Scoring Method#55 source
>>The assumption would be that all of the judges had recieved
>>the same training and are basing their scores on the same
>>range of deductions.
>
>That is an assumption I would never make...
>
>DING DING DING DING DING!!!!!!!!!!
>
>There are many hours spent at the nats to try to get the
>judges on the same page but a simple scan of your sheets
>will show that some people just "Favor" certain things more
>than others.
>
>
>>
>>Let me explain how I judge as an example. I have a very
>>simple system that works for me, and I have a few friends
>>here in Texas that have shared ideas on this method and use
>>it with only minor twists in the deductions from judge to
>>judge.
>>
>>1. The number one objective criteria is shape.
>
>"Without shape there simply is no maneuver"
>
> Tom Farmer.
>
> If the
>>shape is wrong, for several reasons, the deduction counts
>>double over a simple bobble or wiggle. Shape errors are
>>obvious to anyone would looks at the rulebook description so
>>I will not into the detail. A shape error counts two points
>>per occurrance except in the instance of the wingover and
>>the hourglass where they count 4 points per occurrance.
>>This is due to being single manuevers.
>
>So there is a slight change in direction at the top of the
>WO on both tracks that moves him to 32. He also missed int
>on the repeat side of the arc, at the begining, that is
>another 4 to move him to 28. Then he pulls out at 8' on the
>inverted side so he is at 26. ON the upright side he is at
>7' so he is at 25. There was bobble on the outside pullout
>that moves him 24. I have seen this very WO from many top
>pilots around here and they arent getting 24s. I know you
>have gotten a 24 a time or two and not been too happy about
>it. Thing is the WO I described is VERY Common anywhere you
>go. VERY few have even bottoms, no direction changes at the
>top and perfect ints., and no bobbles or very soft inverted
>pullouts. By your standards Bob Gieseke would very easily
>be the high judge...
>
>The clover while
>>being a single manuever has 4 loops and two intersections so
>>a 2 point deduction is sufficient. Each portion of the
>>manuever has its own individual shape error deduction (for
>>example, the square eight has 6 potential incorrect shapes,
>>4 squares and two intersections). Shape errors include any
>>crooked legs, hooks in rounds or flats, size, and gross
>>intersection mistakes.
>
>Dont forget the repeatablity thing...you are supposed to
>repeat the first maneuver you do the second time by. So if
>he blows it the first time but repeats it you are not to
>kill him on that second shape error. Talk to Big Art about
>that. Your method completely ignores this. How come?
>
>>
>>2. A bobble is one point for each occurrance.
>
>Hard corner with a bobble 1 point. Medium with no deviation
>2 points...this rewards the flier trying for RB corners????
>
>>
>>3. The number one subjective criteria is corner size for
>>angled manuevers. There are 3 catagories that I use. Hard
>>(close to rulebook), medium (most everyone), and soft. The
>>flier would recieve an overall dedcution depending on what
>>catagory he falls under. Hard corners receive no deduction,
>>medium two points, and soft 4 points. These are assessed at
>>the completion of the manuever. The total score would be
>>assessed starting at the new potential score. For the
>>starting point would be 40 for hard corners, 38 for medium
>>corners and 36 for soft corners.
>>
>>4. The last deduction catagory for me is bottom height.
>>This is such a common mistake that I have made it a standard
>>deduction as it is easier to keep up with. This is reserved
>>for the fliers that fly good patterns but the bottoms are
>>obviously not at the 4 to 6 foot range. I use a deduction
>>of 1 point per foot. So a 10 foot bottom manuever would
>>receive an automatic 4 point deduction so the flier would be
>>starting a 36 point perfect score.
>
>How do you see a 10 foot or a 9 foot bottom at 150 feet
>away??
>
>Are size errors a deduction in shape as well or do you count
>them seperate? Like size is a deduction and shape is a
>deduction. Some feel too large automatically kills the
>shape. Others say it is simply a separate deduction and
>look at the shape from there. That is why you sometimes see
>a larger pattern get killer scores.
>
>I have looked this over a few times. Heck I couldnt get a
>25 in front of you on any maneuver. Especially my inverted
>flight...
>
>When was the last time you judged using this method? You
>will be the low judge for sure. Is two laps long enough to
>do all the figuring?
>
>I like the fact that you have come up with a method for
>assessing a score. I would like to hear what others use. I
>dont have a method. Your data set would kick me out...
>But then again I dont judge very often. I do feel I have a
>good idea of what it should look like and when it doesnt I
>count down from there.
>
>>
>>Using this system no particular "style" is at a
>>disadvantage. If your style is to try to fly the pattern
>>EXACTLY as stated in the rulebook, go for it, you will be
>>starting at the greatest potential score! You will also be
>>most likely to make errors over another flier who is
>>compromising shape or corners for precision.
>>
>>It is likely however that a compromise flight could score
>>better if the excecution is superb.
>
>Uh what? That totaly contradicts the above paragraph. I
>seem to remember someone really pissed that they tried to
>for RB sizes and a larger pattern beat them. I cant believe
>you would now say that would be ok. FACT IS! If you
>compromise your pattern then your execution can NEVER be
>superb. It is already COMPROMISED! That sentence really
>contradicts itself.

Size is a shape error. It has the largest deduction. No matter how well excecuted the score would not be real good.

>
>Let say you are flying
>>a soft corner pattern, but your excecution of shape and size
>>are extraordinary. With zero bobbles you could achieve a 36
>>for any manuever with corners. That is a great score, but
>>you better be PERFECT on shape and not have any wiggles or
>>the score could drop from the already comprimizing 4 point
>>shortage.
>
>Should only be a 2 point deduction. He repeats the mistake
>on purpose ust as he is supposed to. Then add 1 back in for
>that very good thing that he did correct, the repeating
>thing.
>
>What happend when said flier flies everything really well
>size shape bottoms. However he is really going for RB
>corners and getting bobbles here and there. He should score
>better since he is trying for RB pattern. The compromise
>should never outscore. Dont you think so? Does this happen
>in real life? HECK NO!

Maybe that is why I have devised a system that I think is fair.

>
>That is totaly unfair to the guy busting it to get to the RB
>pattern and trying for it in compitition. This was this
>very fliers argument. He was striving for the RB and others
>were backing off of it and getting ahead. The RB is what we
>all strive for and those backing off should be dinged for
>it.


I think my system would offset the advantage of doing a big, soft pattern. You really should do the math.

>
>>
>>As stated before, size is considered a shape error. It is
>>also deducted at 2 points FOR EACH OCCURRANCE.
>
>Remember the repeat the maneuver thing....
>
> So a flier
>>doing a square eight with perfect medium corners,
>
>How is a medium corner considered perfect?

No bobbles. Medium corners EXCECUTED perfect.

>
> perfect
>>bottoms, intersections, tops and shapes (geometric shape)
>>would receive a 8 point shape deduction and a 2 point corner
>>deduction. A perfect excecution of the manuever under this
>>criteria would receive a 30.
>
>So he gets a 2 point corner deduction and since they are
>"medium" they also contribute to a shape error? I dont see
>where you said he had a shape error. Howcome the 8 point
>deduction?
>

4 squares too large. 8 points.


>But what about the fact that he has repeated the same thing
>over and over? That has some weight to his advantage. The
>maneuver you have described would probably score ALOT higher
>than 30 on most circles. If medium corners are the only
>issue I will take it every time!
>
>>
>>****I am still undecided as to whether a 2 point deduction
>>for shape for size is fair, I guess it depends on the camp
>>of judges you are talking too, a 1 point deduction per
>>occurance might be a better rule. As long as it is
>>consistance I would not care.*****
>
>Interesting...
>
>I like that you are showing your method. Shows some very
>forward thinking. All judges have their OWN methods right?
>So how in the world can we begin to think they should all
>fall under the same data set and it would somehow be a fair
>assessment of the judges ability?

Because Doug THEY DO!!!! Judges should be measured the same as the fliers, period. If the argument is that all the judges use their own methods, and that is just fine, how in the world can you determine if the judge is not judging on the color of the airplane or some redicualous criteria.

The entire idea that the measure of flying is 100% "objective" (defined by the rulebook) but the measure of the judge measuring the objective is %100 "subjective" makes me crazy.

Absolutely, we should be comparing methods between judges and aligning them. Are you kidding? I posted mine. I am sure there are others that do the same thing and would ACHIEVE RESULTS THAT WOULD COORELATE.

DMoon · Oct 18, 2004 02:33 PM

RE: Scoring Method#65 source
>>So there is a slight change in direction at the top of the
>>WO on both tracks that moves him to 32. He also missed int
>>on the repeat side of the arc, at the begining, that is
>>another 4 to move him to 28. Then he pulls out at 8' on the
>>inverted side so he is at 26. ON the upright side he is at
>>7' so he is at 25. There was bobble on the outside pullout
>>that moves him 24. I have seen this very WO from many top
>>pilots around here and they arent getting 24s. I know you
>>have gotten a 24 a time or two and not been too happy about
>>it. Thing is the WO I described is VERY Common anywhere you
>>go. VERY few have even bottoms, no direction changes at the
>>top and perfect ints., and no bobbles or very soft inverted
>>pullouts. By your standards Bob Gieseke would very easily
>>be the high judge...

YOu siad I should do the math. Well right here above I did do the math. I thought of a WO I see from a pilot around here is always in the top 3 at any contest he enters and this is his score accordig to YOUR math. Next time I get to fly again, who knows when, I will ask you to score his WO as see what you get. I know the judges around here dont judge it near that hard. I would also like to see score sheet of yours that uses this method. Heck score me and lets see what happens. If I ever get to fly agian.

>>
>>Dont forget the repeatablity thing...you are supposed to
>>repeat the first maneuver you do the second time by. So if
>>he blows it the first time but repeats it you are not to
>>kill him on that second shape error. Talk to Big Art about
>>that. Your method completely ignores this. How come?

>>>Using this system no particular "style" is at a
>>>disadvantage. If your style is to try to fly the pattern
>>>EXACTLY as stated in the rulebook, go for it, you will be
>>>starting at the greatest potential score! You will also be
>>>most likely to make errors over another flier who is
>>>compromising shape or corners for precision.
>>>
>>>It is likely however that a compromise flight could score
>>>better if the excecution is superb.
>>
>>Uh what? That totaly contradicts the above paragraph. I
>>seem to remember someone really pissed that they tried to
>>for RB sizes and a larger pattern beat them. I cant believe
>>you would now say that would be ok. FACT IS! If you
>>compromise your pattern then your execution can NEVER be
>>superb. It is already COMPROMISED! That sentence really
>>contradicts itself.
>
>Size is a shape error. It has the largest deduction. No
>matter how well excecuted the score would not be real good.

You never really said that anywhere in your method until now. Thanks for the clarification.

>>Should only be a 2 point deduction. He repeats the mistake
>>on purpose ust as he is supposed to. Then add 1 back in for
>>that very good thing that he did correct, the repeating
>>thing.
>>
>>What happend when said flier flies everything really well
>>size shape bottoms. However he is really going for RB
>>corners and getting bobbles here and there. He should score
>>better since he is trying for RB pattern. The compromise
>>should never outscore. Dont you think so? Does this happen
>>in real life? HECK NO!
>
>Maybe that is why I have devised a system that I think is
>fair.
>
>>
>>That is totaly unfair to the guy busting it to get to the RB
>>pattern and trying for it in compitition. This was this
>>very fliers argument. He was striving for the RB and others
>>were backing off of it and getting ahead. The RB is what we
>>all strive for and those backing off should be dinged for
>>it.
>
>
>I think my system would offset the advantage of doing a big,
>soft pattern. You really should do the math.

I have done the math. See above before you assume.

Also what do you do when you see a NORMAL flight? Meaning there are all types of corners throughout the pattern. How does those fit in?

>
>>
>>>
>>>As stated before, size is considered a shape error. It is
>>>also deducted at 2 points FOR EACH OCCURRANCE.
>>
>>Remember the repeat the maneuver thing....

Rmember when we talked about inverted flight. You totaly stumped me. Dont forget you must use the RB to get your criteria and it says that you need to repeat the first maneuver the second and third time. So you cant kill him for EACH occurance. You nail him on the first time but then he repeats it to the T and you say good job for the precise flying.

>>
>> So a flier
>>>doing a square eight with perfect medium corners,
>>
>>How is a medium corner considered perfect?
>
>No bobbles. Medium corners EXCECUTED perfect.

The medium corner "perfect" no bobbles whatever has no merit in the RB. The book wants us to go for 5' all the time. I dont get that at all.

>
>>
>> perfect
>>>bottoms, intersections, tops and shapes (geometric shape)
>>>would receive a 8 point shape deduction and a 2 point corner
>>>deduction. A perfect excecution of the manuever under this
>>>criteria would receive a 30.
>>
>>So he gets a 2 point corner deduction and since they are
>>"medium" they also contribute to a shape error? I dont see
>>where you said he had a shape error. Howcome the 8 point
>>deduction?
>>
>
>4 squares too large. 8 points.

Dont forget repeating the maneuver. Should only be counted down on the first 8 and not on the second 8.

>>Interesting...
>>
>>I like that you are showing your method. Shows some very
>>forward thinking. All judges have their OWN methods right?
>>So how in the world can we begin to think they should all
>>fall under the same data set and it would somehow be a fair
>>assessment of the judges ability?
>
>Because Doug THEY DO!!!! Judges should be measured the same
>as the fliers, period. If the argument is that all the
>judges use their own methods, and that is just fine, how in
>the world can you determine if the judge is not judging on
>the color of the airplane or some redicualous criteria.

That is just it. YOU CANT! How are you going to judge the judges when the book gives no specifics on how to have a certain criteria, like the one you laid out as an example.

I guess we need to judge the judges? Who is going to judge them...

>
>The entire idea that the measure of flying is 100%
>"objective" (defined by the rulebook) but the measure of the
>judge measuring the objective is %100 "subjective" makes me
>crazy.

Sure it does. Me too.

>
>Absolutely, we should be comparing methods between judges
>and aligning them. Are you kidding?

Ok aligning the like methods yes. But compering unlike methods and saying one is wrong and the other is right. This can only be done when a judge uses NON RB criteria, such as me and inverted flight. And no I am not kidding.

But as I showed in my first post two judges using RB criteria cancome up with vastly different scores on the same maneuver. I contend this will never change.

I posted mine. I am
>sure there are others that do the same thing and would
>ACHIEVE RESULTS THAT WOULD COORELATE.

Maybe and maybe not.

50plusAirYears · Oct 18, 2004 04:11 PM

RE: Scoring Method#73 source
Is it time yet to send the referee home, take off the gloves, and really get down and dirty?

Brett Buck · Oct 18, 2004 04:31 PM

RE: Scoring Method#74 source
>>Because Doug THEY DO!!!! Judges should be measured the same
>>as the fliers, period. If the argument is that all the
>>judges use their own methods, and that is just fine, how in
>>the world can you determine if the judge is not judging on
>>the color of the airplane or some redicualous criteria.
>
>That is just it. YOU CANT! How are you going to judge the
>judges when the book gives no specifics on how to have a
>certain criteria, like the one you laid out as an example.

Exactly on point.

>>The entire idea that the measure of flying is 100%
>>"objective" (defined by the rulebook) but the measure of the
>>judge measuring the objective is %100 "subjective" makes me
>>crazy.
>
>Sure it does. Me too.
>

But any concept that the measure of flying is "objective" is completely incorrect. It is entirely *subjective*. There is set of defined perfect maneuvers. But there is no standards specified for deviations from perfect (aside from a few simple things like "falls outside the range of 3.9-5.9 feet which are wholly inadequate to the task). Even if such items did exist, there is nothing remotely resembling an objective method for measuring deviations.

Someone eyeballs a squiggle in the air, and comes up with a score for it. That's it. They are not supposed to think, "well, it's Paul Walker so that must have been a 38", or "I hate green airplanes" but it's a long way from avoiding that sort of reasoning, to something resembling objectivity.

The goal may well be complete objectivity, but there has been no practical system proposed to make this happen. Any idea that it's objective now, and that the rules and available techniques define anything resembling objectivity, is (sorry Brad) delusional.

That's why this idea of "filtering" or "detecting" errant judges is invalid at it's root. It makes absolutely no difference which equations one uses - the underlying premise that there is some objectively "right" score is simply faulty. If it's subjective, then there are no right or wrong answers. So it's patently impossible to come up with an objective system to root out errant answers.

It's a faulty idea because the premise is wrong.

Brett

godzilla · Oct 18, 2004 05:30 PM

RE: Scoring Methodedited#79 source
>
> It's a faulty idea because the premise is wrong.
>
>Brett

OK Brett, I want to make nice.

I appoligize for implying that Phil might dot your eye because you said he was dillusional. So there, I am sorry.

I do respect your big ol' brain. I have told you so and have told several others the same, from Coast to Coast. I would like to think you share my sentiment, but apparently this is not true.

I read this post. I understand what you are saying.

Now give me a chance here:

I gave my statistical resume. I am a rank beginner, I do however have some good tools.

The thing that appears to be misleading is that the data cannot be studied using STANDARD statistical tools. It could be studied to prove your assessment that it is %100 objective, as you said. You could assert that this might be true today. I would however like to see a study to see if it were true. I do believe, in my heart, that judges if properly trained, and with sufficient experience track each other. This coorelation be shown and efforts made to improve this.

I think in the future steps could be taken to improve the coorelation by improving the objective portion of the scoring. True, I know of no universal scoring system, obviously none exists today. That is not to say that steps could not be taken to improve. This is Process Improvement 101. I posted my system, not to say it should be adopted, but to say that it tended to show good coorelation from judge to judge in the small tests I have made.

Perhaps it would be good to look at putting steps put in place to increase coorelation over the coming years possible if it were found that the coorelation were sufficiently low. I do believe that a ratio of objective to subjective judging criteria would be universally good for all.

If you disagree, fine. You stated your opinion. It is however, one opinion. It is obvious that not everyone completely agrees with you.

No method has been stated for anything I have done. I have no burden of proof because I have made no statement that I have a proof. I do however have tools that are specifically designed to study data EXACTLY LIKE JUDGING.

In Six Sigma you have process inputs. This includes the who, with what, etc that effect the process. From what you are saying, it appears that you feel there insufficient inputs into the process.

For example, the process inputs could include:
1. Rulebook.
2. Judge.
3. Training.
4. Flyer.
5. Manuever excecution.

All effect the outputs of the process which would be:
1. A score.

To be more objective a new input would have to added, a METHOD.

So, based on that fact, you might have a point that my play data has no basis on what is done today. That is fine.

50plusAirYears · Oct 18, 2004 05:35 PM

RE: Scoring Method#81 source
"Myself when young did eagerly frequent Doctor and Saint,
and heard great arguement about it and about,
And evermore,
Went out the same door wherein I came!"

Sometime about 1958, a couple Stunt pilots tried to address this issue by constructing a "Judging Aid". They built a wire framework contraption over a bowling ball Showing to scale what the "Ideal Pattern" would look like to a judge sitting the prescribed distance from the circle. Of course, the same pattern looked different to the judges sitting on each side of the center judge. If the competitor didn't have the manuever centered in the center judge, an other wise perfect manuever would look flawed. It never became popular. It was an attempt to introduce at least some degree of objectivity to judging by giving a template the judges could compare to. Without something like this the judges have to guess at the positioning, corner radius, 45 degree angles, etc. on a three dimension body of air.

I've even read treatises stating the 5 foot radius is impossible to achieve with any kind of stunt ship. Lots of math to back it up, and even measurements done using time-lapse photography.

Someone else did some night pattern flying for Flying Models using a pen light strapped to the outside wing and a camera with a long open time. A perfect loop from the pilot's point of view at the center of the circle was shown by the camera outside the circle to be unrecognizeable as a loop. But, a set of judges might have scored it near perfect.

Judging by eyeball is at least 95% subjective, maybe 5% objective. To be more objective, the judges need some kind of measuring system that gives them a real time temolate to compare against.

Probably nobody has yet come up with an acceptable template.

The original proposition sounds like a fun project for Green or Black Belt practice.

Now how do i unsubscribe to this thread. I'm getting too many reply notifications at work here.

ferocious · Oct 18, 2004 05:48 PM

RE: Scoring Method#82 source
>....Someone eyeballs a squiggle in the air, and comes up
>with a score for it. That's it.....
> The goal may well be complete objectivity, but there has
>been no practical system proposed to make this happen. Any
>idea that it's objective now, and that the rules and
>available techniques define anything resembling objectivity,
>is (sorry Brad) delusional.

One judge is an opinion. Two is an argument, three is a disagreement. 25 judges and it is a fact. Use more judges. Problem solved.

> It's a faulty idea because the premise is wrong.
>
>Brett


Ain't it a shame when technology catches up with long-held opinions. Thanks to Keith Renecle, we now have an object demonstration of the perfect pattern, animated on the computer screen. Anyone can now see exactly how a correct pattern should look, from almost any point in and around the circle.

we still need 25 judges though.

grzly23 · Oct 18, 2004 06:28 PM

RE: Scoring Method#83 source
Phil,

Correct according to who? Keith or the rule book?


>>....Someone eyeballs a squiggle in the air, and comes up
>>with a score for it. That's it.....
>> The goal may well be complete objectivity, but there has
>>been no practical system proposed to make this happen. Any
>>idea that it's objective now, and that the rules and
>>available techniques define anything resembling objectivity,
>>is (sorry Brad) delusional.
>
>One judge is an opinion. Two is an argument, three is a
>disagreement. 25 judges and it is a fact. Use more
>judges. Problem solved.
>
>> It's a faulty idea because the premise is wrong.
>>
>>Brett
>
>
>Ain't it a shame when technology catches up with long-held
>opinions. Thanks to Keith Renecle, we now have an object
>demonstration of the perfect pattern, animated on the
>computer screen. Anyone can now see exactly how a correct
>pattern should look, from almost any point in and around the
>circle.
>
>we still need 25 judges though.

godzilla · Oct 18, 2004 06:50 PM

RE: Scoring Method#84 source
>Phil,
>
>Correct according to who? Keith or the rule book?
>

The Rulebook Griz!!!!

The rulebook is simply geometry. The simulation is a moving example of the diagrams used today. The difference is that the airplane is moving, giving one a sense of reality and not just shapes in a book.

It is really awesome. Really. You should check it out.

Keith Renecle · Oct 20, 2004 12:51 PM

RE: Scoring Method#87 source
Hi All,
This is another one of those “juicy” threads. Phil, thanks for the support on the simulation, and to “Griz” I don’t take any comments personally about the sim. Your question about “Correct according to who? Keith or the rule book?”, is a very valid one. It is also one that has been asked since I started on trying to create some way of teaching the perspective of all of the maneuvers to our local judges. I thought that a neat way to demonstrate this would be to create a simulation something like the R/C sims, except that the pattern should be flown automatically and not by the viewer. This was most difficult for me, because I am not a computer programmer. I found that I could use 3-D game development software, but it was still a difficult learning curve for me. The sim right now is still very basic and has some glitches, but it gets better every day. I also get the boffins like Pete Soule and Igor to check my work. Igor regularly gets me to change things!

To make sure that I get everything right, I draw a wire-frame hemisphere, and then accurately draw each maneuver shape on the surface, following the rule definitions. My 3-D model airplane, rotates from the centre of the sphere, and by rotating each axis, I step the model over the shape, frame by frame. Right now I have stopped at the clover, and it is over 8000 frames. This is something like doing a Disney animation. The big difference to a normal animated movie, is that you can move the various camera’s (viewpoints) to anywhere in the virtual world around the flying circle. This creates an incredible opportunity to watch a perfect (almost!!) “to-the-rules” (FAI) pattern that you can watch over and over again, and from any angle or viewpoint. I have follow-camera’s and fixed camera’s, and just recently for my own fun, I stuck an extra camera on the model itself. This is wild!

I said that I stopped at the clover, because I just can’t draw any of the shapes from the AMA or FAI rules, that will fit the sphere 100%. If you saw Igor’s thread with my drawings on the clover, you may just understand the problem. There are a couple of ways to “fudge” the shape, but strictly speaking, it just cannot be done to the rules. Igor is helping me at the moment to try and get the rule makers to come to some conclusion. I also created a simulation of three possible ways to perform the clover, with pop-up shapes to illustrate the differences, and I’m willing to share this with anyone that is interested. I have not advertised my sim as yet, because it is still so rough, and it is unfinished. As soon as it is completed, I will stick it on a website for free download. Quite a few stunt folks have copies already, and so far, I have only had good comments.

I agree with most of Brett’s comments about stats. They are most interesting, and I always study Bruno Delor’s good spreadsheets after each world champs. The big thing is obviously that we would love to know who is judging closest to the rules, and not to some norm. Without some kind of “feed-back” from the model, this will never be possible. BTW has anyone in C/L ever looked at the TBL system used in R/C pattern? I have a copy of the principles of the newer TBLP system, but it’s not software, just a paper copy. I’m pretty sure that it must be available somewhere on the net. They use around 10 judges, and some rather fancy “fuzzy” logic to normalize they scores. We will never get rid of the human problems in stunt judging, but we can improve judging drastically, by getting international consensus on how the shapes should look. This is not going to be easy, but it is possible. Good, positive debates, by open-minded, experienced people can only benefit all of us.

grzly23 · Oct 20, 2004 02:13 PM

RE: Scoring Method#88 source
Keith,

Thanks for taking time to answer my question. I am grateful you interpreted the spirit of my question and that it was not to be contentious or argumentative. Your simulation effort is truly outstanding and worthy of praise. Don't give up. What you have proven so far is that reality is so much more difficult than simulation. But, I know you will persist and succeed. What turn radius are you using? The one called for in the rule book of 5 feet or the actual radius flown by the majority of stunt ships? I hope it is the latter. If and when you succeed, it should give all of us a wonderful judging, practice and training tool to improve our flying and judging capabilities.

godzilla · Oct 20, 2004 03:49 PM

RE: Scoring Method#90 source
>Keith,
>
>Thanks for taking time to answer my question. I am grateful
>you interpreted the spirit of my question and that it was
>not to be contentious or argumentative. Your simulation
>effort is truly outstanding and worthy of praise. Don't
>give up. What you have proven so far is that reality is so
>much more difficult than simulation. But, I know you will
>persist and succeed. What turn radius are you using? The
>one called for in the rule book of 5 feet or the actual
>radius flown by the majority of stunt ships? I hope it is
>the latter. If and when you succeed, it should give all of
>us a wonderful judging, practice and training tool to
>improve our flying and judging capabilities.

The radius is 5 feet.

Here, here! One everything else. We will all be in this man's debt.

david eyskens · Oct 20, 2004 03:57 PM

RE: Scoring Method#91 source
I am very curios, I do not know you, knowbody's fault....You have now been handed the complete and total responsibility of managing the judging process for the Nats 2005--It is your GIG totally...In a nustshell what would you do with the responsibility??Thanks David

godzilla · Oct 20, 2004 05:56 PM

RE: Scoring Method#95 source
>I am very curios, I do not know you, knowbody's fault....You
>have now been handed the complete and total responsibility
>of managing the judging process for the Nats 2005--It is
>your GIG totally...In a nustshell what would you do with the
>responsibility??Thanks David

Not a lot at first. I think that would be irresponsible. Rome was not built in a day.

1. I would post a giant add in the PAMPA mag asking for judges for the 2005 Nats. Any responses would be shared with the district rep for PAMPA.
2. I would use feedback gathered by the district rep to establish who might be the most qualified new candidates. I would then start these candidates on sim training.
3. I would ask for scores from the last two Nats. I would attempt to establish a baseline based on the current data.
4. I would suggest that the current judges with the the lowest coorelations in the past should be contacted by their district rep or myself to establish their philosophy for judging. In fact, it would be a good idea to get feedback from all the current judges on how they score (to the best of their abilities). I would also ask for any wholistic views on what they are seeing in the current crop of patterns (as Peabody did on the hourglass). It would be incredibly useful to know "pet" scoring criteria from the entire crop of judges (shape over intersections over corner etc)
5. I would implement sim training for the judges training ASAP. I would suggest distributing the sim to every judge immediately and then gather feedback on what they see.
6. I would begin to formulate ideas for standard deductions and get feedback from all current judges.
7. I would report in Stunt News what I am doing.
8. Based on the entire process I would propose rules changes in the future to attempt to increase coorelation.

In a nutshell, I guess I would try to open some lines of communication to the entire judging and CLPA community.

david eyskens · Oct 21, 2004 12:41 PM

RE: Scoring Method#104 source
Godzilla,
First and foremost, I want to express my respect for your passion, & enthuisiam regarding this issue..It is evident that this issue is important to you....This event should never sit in "concrete" and state: this is the way it is going to be "forever"!!!This is also an event that historically has been in my opinion a bit "conservative"
Attempting to initiate change, can be extremely difficult, and you will experience the the effect of such...Harness your passion constructively, and allow the process to go where it may...I wrote an artical in Stunt News in the late 90's, my observation and critical thinking centered around the "pyschology" of judging..Meaning that the subjective nature of judging takes on a very complex pyschological process, and if the person assigned such a difficult responsibility as judging, is not aware of the intrapersonal aspect of this process, serious problems will occur.....I am convinced that it is the "pyschological" intra-psyche aspect of judging that has been ignored, and where the root cause of the issue resides..Until this is recognized and addressed the process will remain where it is....Be aware I am not saying that it is bad where it stands now.Thanks David

godzilla · Oct 21, 2004 01:44 PM

RE: Scoring Method#105 source
....Be aware I am not
>saying that it is bad where it stands now.Thanks David

Neither am I.

Where are we now? I would really like to know. To know where you are in a process you must have measures and you must measure with tools. All I would do is start the process of measurement.

Brett Buck · Oct 20, 2004 02:35 PM

RE: Scoring Method#89 source
>>....Someone eyeballs a squiggle in the air, and comes up
>>with a score for it. That's it.....
>> The goal may well be complete objectivity, but there has
>>been no practical system proposed to make this happen. Any
>>idea that it's objective now, and that the rules and
>>available techniques define anything resembling objectivity,
>>is (sorry Brad) delusional.
>
>One judge is an opinion. Two is an argument, three is a
>disagreement. 25 judges and it is a fact. Use more
>judges. Problem solved.
>
>> It's a faulty idea because the premise is wrong.
>>
>>Brett
>
>
>Ain't it a shame when technology catches up with long-held
>opinions. Thanks to Keith Renecle, we now have an object
>demonstration of the perfect pattern, animated on the
>computer screen. Anyone can now see exactly how a correct
>pattern should look, from almost any point in and around the
>circle.


True, such a program exists. But that doesn't have anything to do with the point I was making. We already had the definition of a perfect maneuver written into the rulebook. Keith's program just makes it graphical. That has nothing to do with the fact that assessing the errors with respect to the perfect maneuver is completely subjective. Both in terms of defining what short and magnitude of errors get particular deductions, and in the ability to determine the magnitude and nature of the errors themselves.


>we still need 25 judges though.

I once again note that 2,5,25 notwithstanding, you are misapplying statistics. The lack of sample size was just a quickie way to prove it. But even if you had 100 judges, the variation between them *still means nothing* in a statistical sense. More judges will make your calculated standard deviation go down, but that still ignores the inconvenient fact that at the root of the issue is you still consider a particular score objectively "right", and that deviations are the result of random errors with respect to this "right" score. Neither of which is a correct assumption, and it can never be unless you remove the subjectivity as described above. It's not random errors if it's based on something, which in this case it most certainly is. And if it's not random errors, then the statistical measures you are trying to use *don't apply*.

That's not a matter of my opinion - it's mathematically incorrect, as you are violating the basic assumptions under which the measures were derived in the first place. I don't know how much more clear it can be.

Brett

ferocious · Oct 20, 2004 06:17 PM

RE: Scoring Method#96 source
Well Brett, you might want to take on a second career as a statistics teacher. There are thousands of statisticians in industry who are using statistics on judging-type data everyday. It seems to work quite well for them, but you could enlighten them to the error of their ways.

The science of statistics really doesn't care how you got the numbers. You can generate them by flipping coins, throwing dice, or judging stunt patterns. I agree though that you can't, from the judging data, work backwards and really tell HOW the judges are getting from the stimulus(the maneuvers) to the score. Each judge may have a totally different way of doing this, which really doesn't matter, as long as they can consistently produce results that can differentiate between the flyers.

Brad has a very good point about judging methods. If a judges can't go back to the pattern just flown and give a reasonable explanation of how they arrived at scores his score is kind of a problem. The rulebook gives some criteria that they should be paying attention to.

Brett Buck · Oct 20, 2004 11:34 PM

RE: Scoring Method#99 source
>Well Brett, you might want to take on a second career as a
>statistics teacher. There are thousands of statisticians in
>industry who are using statistics on judging-type data
>everyday

Which neatly avoids addressing any of my counter-points. You failed to address them the previous three times, too, and also degenerated to sarcasm there, too. I'm perfectly willing to argue real points with you, assuming you have some.

And only since you mentioned this the previous time, if you want to impress everone with either dollar value or risk to human life based these experts - I'll almost certainly win that one, too.


> It seems to work quite well for them, but you
>could enlighten them to the error of their ways.


Actually, funny you should mention that, but I did teach a statistics course in a professional setting. After demolishing the Six Sigma "expert" when their cult revival meeting came to town, and with the enthusiastic support of the numerous aerospace engineering PhD's in the audience, the guy finally admitted it didn't apply in our case and that several of his assuptions were invalid in our situation (although for completely different reasons than here, although six-sigma process controls were also inapplicable for the N too small reason, too) and our management was so impressed that I got to/was induced to by threat, give a series of lectures on the topic.

Didn't really take, of course - went over with real engineers just fine, but it made it completely obvious that the "quality engineers" were just there to keep someone from stealing their desks, and to emit 65 watts of heat to help with our winter energy costs.


A digression - but on that topic - the underlying concepts of Six Sigma are sound as they can be, but in acual application it is frequently treated far more like a religion than an industrial process control concept. It's valid if you are producing 6 million bullets a month - but it doesn't work so good at all when you are talking about manufacturing 6 custom-built satellites over the period of 15 years.

Then along came ISO 9xxx- leading to proud moments like :

http://www.spaceflightnow.com/news/n0410/04noaanreport/

The result of independent thought being replaced with "process" - which they followed to the letter. It said "wipe ring with IPA -soaked cloth" - not "wipe ring with IPA-soaked cloth and at the same time notice that the bolts were missing"

Brett

Dave Simons · Oct 20, 2004 11:49 PM

RE: Scoring Method#100 source
"Then along came ISO9XX..."

Please!decorum! - ISO9xxx is not mentioned in polite company, around here anyway.

Only thing I've ever seen ISO 9etc do is assign blame and replace engineers with MBA's. Engineering organisations usually dont last too long after that.

50plusAirYears · Oct 21, 2004 08:47 AM

RE: Scoring Method#102 source
>"Then along came ISO9XX..."
>
>Please!decorum! - ISO9xxx is not mentioned in polite
>company, around here anyway.
>
>Only thing I've ever seen ISO 9etc do is assign blame and
>replace engineers with MBA's. Engineering organisations
>usually dont last too long after that.

We got lucky on that. Our former parent company MBA CEO sold us off to a German company. Their CEO has Masters and Doctorates in both Mechanical and ElectronicEngineering. Our former CEO grew us from 650 million in sales to about 520 million, reduced our Engineering department from 120 to 98, increased the bean counters from about 40 to 60, but we were always "Making the Numbers". Lost us Teir One status with most of our customers and even lost us a couple important customers entirely. In less than 2 years, we are now looking to hit 600 million this year, we have regained lost customers, We are now the preferred supplier to most of them, our Engineering dept now is at 125, and projected to grow over the next couple of years. We now are focused on developing and selling customer satisfying products. ISO 9000, 9100, 14000, etc, and Six Sigma and the like are now tools to be applied where useful, not "Our Corporate Culture".
The former CEO did raise the Corporate margin from about 6% to about 6.08% in 12 years, but we were running out of Divisions to sell off.

godzilla · Oct 22, 2004 03:47 AM

RE: Scoring Method#112 source
> Actually, funny you should mention that, but I did teach
>a statistics course in a professional setting. After
>demolishing the Six Sigma "expert" when their cult revival
>meeting came to town,

> Didn't really take, of course - went over with real
>engineers just fine, but it made it completely obvious that
>the "quality engineers" were just there to keep someone from
>stealing their desks, and to emit 65 watts of heat to help
>with our winter energy costs.

Wow. No wonder.

> A digression - but on that topic - the underlying
>concepts of Six Sigma are sound as they can be, but in acual
>application it is frequently treated far more like a
>religion than an industrial process control concept.

Six Sigma is NOT A "PROCESS CONTROL" CONCEPT. Never has been, never will be.

Another case of you being misinformed, or misinforming I cannot tell which.

ISO is a process control tool. Six Sigma is not. Six Sigma has process controls in its method (none originating with Six Sigma BTW), but it is not a process control philosophy. Whoever implied this to you was also misinformed. You may be referrring to CTQ, which is really more of a buzz word:

"Critical To Quality - CTQ

CTQs (Critical to Quality) are the key measurable characteristics of a product or process whose performance standards or specification limits must be met in order to satisfy the customer. They align improvement or design efforts with customer requirements.

CTQs represent the product or service characteristics that are defined by the customer (internal or external). They may include the upper and lower specification limits or any other factors related to the product or service. A CTQ usually must be interpreted from a qualitative customer statement to an actionable, quantitative business specification.

To put it in layman's terms, CTQs are what the customer expects of a product... the spoken needs of the customer. The customer may often express this in plain English, but it is up to us to convert them to measurable terms using tools such as DFMEA, etc."

Six Sigma is mainly a problem solving method.

Six Sigma

The goal of Six Sigma is to increase profits by eliminating variability, defects and waste that undermine customer loyalty.

"Six Sigma can be understood/perceived at three levels:

1. Metric: 3.4 Defects Per Million Opportunities. DPMO allows you to take complexity of product/process into account. Rule of thumb is to consider at least three opportunities for a physical part/component - one for form, one for fit and one for function, in absence of better considerations. Also you want to be Six Sigma in the Critical to Quality characteristics and not the whole unit/characteristics.
2. Methodology: DMAIC/DFSS structured problem solving roadmap and tools.
3. Philosophy: Reduce variation in your business and take customer-focused, data driven decisions. "

If you were more informed I think you could find the problem solving methods useful for solving hard to solve problems in a group.

>valid if you are producing 6 million bullets a month - but
>it doesn't work so good at all when you are talking about
>manufacturing 6 custom-built satellites over the period of
>15 years.

Once again, a misstatement. Six Sigma's DMAIC roadmap would be PERFECT for a custom "one off" shop in certain cases, but I do not think that is what is was designed for. Six Sigma is really, at its heart, a method by which MONEY can be found hiding in less than obvious spots. Money is important to companies that have to COMPETE and are held financially accountable by their investors and stockholders. These companies have to SHOW A PROFIT to stay in business. Something that has not really been a high priority in the space business.

So, no I would think that Six Sigma would be an odd match, since most of the space business is more akin to Socialism than Capitalism.

>
>Then along came ISO 9xxx- leading to proud moments like :

>http://www.spaceflightnow.com/news/n0410/04noaanreport/

>The result of independent thought being replaced with "process" - >which they followed to the letter. It said "wipe ring with IPA >-soaked cloth" - not "wipe ring with IPA-soaked cloth and at the same >time notice that the bolts were missing"
-------------------------------------------

I have very little use for ISO, it is mainly bookkeeping, but something like it must be done, as process documentation is critical to maintain consistency. I find it very interesting that you are implying that process control was the cause of the mishap and not the FAILURE to have an accurate process control.

That would be simple incompetence.

What would be the alternative to process controls? No work instructions?

klelmore · Oct 20, 2004 10:21 PM

RE: Scoring Methodedited#98 source
>>we still need 25 judges though.
>
> I once again note that 2,5,25 notwithstanding, you are
>misapplying statistics. The lack of sample size was just a
>quickie way to prove it. But even if you had 100 judges,
>the variation between them *still means nothing* in a
>statistical sense. More judges will make your calculated
>standard deviation go down, but that still ignores the
>inconvenient fact that at the root of the issue is you still
>consider a particular score objectively "right", and that
>deviations are the result of random errors with respect to
>this "right" score. Neither of which is a correct
>assumption, and it can never be unless you remove the
>subjectivity as described above. It's not random errors if
>it's based on something, which in this case it most
>certainly is. And if it's not random errors, then the
>statistical measures you are trying to use *don't apply*.
>
> That's not a matter of my opinion - it's mathematically
>incorrect, as you are violating the basic assumptions under
>which the measures were derived in the first place. I don't
>know how much more clear it can be.
>
> Brett

I'll add this and then, I promise, I'll shut up. Brett's assertion that violating underlying assumptions makes the method (whatever method it is) invalid, while technically correct, isn't always practical. We *never* satisfy all our underlying assumptions when we do parametric statistics. In real life, we almost never have access to the entire "population" (whatever that really is), and our sample is almost never truly random. Our data is usually whatever we can get. Error sources are almost never really random in that there is nearly always an underlying (but unknown) process at work. If the error source looks "randome enough," however, we can still play.

The failure to meet our assumptions boils down to questions about robustness: is the statistic used robust when the underlying assumptions are not satisfied? How "gracefully" do the statistical parameters fail as we stray further and further from the assumptions? In what way do the statistics fail? More Type I errors or more Type II errors? Or both? The answer is almost never straightforward and seldom a simple yes/no.

Remember, however an excellent point Brett made earlier: the numerical scores are not as important as the *ranking.* Rank is everything in a contest: we eant to know who is ranked highest because that pilot gets the trophy. We simply use the numerical scores as a way to define rank. As long as the judges act consistently thoughout a contest, we may expect the rankings to remain robust and valid, even if one judge is generating completely random, uncorrelated scores throughout the event.

Kim Elmore

godzilla · Oct 20, 2004 05:37 PM

RE: Scoring Method#92 source

> The goal may well be complete objectivity, but there has
>been no practical system proposed to make this happen. Any
>idea that it's objective now, and that the rules and
>available techniques define anything resembling objectivity,
>is (sorry Brad) delusional.

YES! YES!!! TRUE!!!!

I never said there was current system for objective scoring!!!! I said there was an objective guide used to make a subjective judgement!!! You are so correct... but it does not change a thing. the assumption that well trained, experienced judges will vascilate at a cooresponding rate is perfectly reasonable. Also, the assertion that have %100 coorelation would be a goal for quality judging is also reasonable.

As I have said more than once, the idea was to study the coorelation, be it subjective or objective really matters not. It is perfectly reasonable to study coorelations, be the measure subjective or objective.

In fact, it would be my contention if that we did study current data, it would be used a baseline only. Everyone might be surprised at the high coorelation of scores today. If, in the future, we did implement changes to try to improve the judges training (for example) we could measure the success of this process (or failure for that matter).

As to the mathmatical incorrectness of the assumptions of "the puzzle". I have no other way of saying it, as far as I can tell with the statements you have applied to your debate either you are flat wrong, misinformed, or are picking on some samantics that no one appears to understand (kind of like every time someone says the word "power" or "thrust" you have to explain to them how they are wrong, but yet you have proposed no other word that would be correct to describe the feeling or "more power").

In my phone conversation with Phil C. he described studies that were made at Hershey (the candy bar company). Phil recently retired from there after a full career (despite what you may think Phil is no "poser" on this subject). Apparently a lot of his time spent at Hershey was STUDYING SUBJECTIVE JUDGEMENTS BASED ON OBJECTIVE CRITERIA. It is called TASTE TESTING. In a taste test, large amounts of data is collected based on completely subjective data using objective guidelines. It is still necessary to determine the less valid parts of the data sets.

Think about it. Have you ever completed a taste survey? I have for Frito Lay. Usually there is some type of scale for taste or some other characteristic, an explanation of how to determine the score, and then it is completely up to the testee to come up with a score. As you know, taste is subjective to infinite degree.

Sound familiar?

I can give you several other examples of measuring and comparing of subjective scores used in industry (customer care surveys for example). The idea that values in data that are subjective are not good data and cannot be measured is simply an ideal not shared by the rest of the world. Also, the idea that PEOPLE DO NOT MAKE BILLION DOLLAR DECISIONS BASED ON THIS DATA WOULD ALSO BE WRONG. People measure and compare subjective data every day and all over the the world, AND THEY USE STATISTICAL MEASURES AND TESTS TO DO IT. Maybe not in the rocket ship business, but there are other industries.

...and what a cooincidence. The test I found to be significant for the "puzzle" turned out to be one of the exact test used at Hershey. Not a cooincidence, a common method!!!

You keep bringing up N=3, I assume you are talking about the matrix size. I do not believe this applies, or somehow you are misapplying the rule you think you are applying (since you have not stated the rule). The minimum number of judges required to test coorelation would be two, which is one more than one. The number of fliers needed would be one. The data set would really only need to be about 30 scores for a p value > .05. Why use .05? It is standard.

The assumed hypothesis (Ha)= there is a coorelation between the judges scores.
The null hypothesis (Ho) = there is no coorelation between the judges scores.

Are you saying there are a significant number of data points to have a valid test because one axis of the matrix is only 4? I have no idea where you got this idea, but I am sure it is misapplied. Are saying that no matter how many actual "tests" we have for comparison, the only criteria for a good test is the matrix size?

....and truly Brett, I am not saying any of these things to badger or belittle you, I believe the only explanantion is that there is some communication gap, same as what I said to Ted's post. I do believe I do understand your concerns, both of yours, but I truly believe you do not undersatand what I and others who have expressed their opinions are trying to do or maybe how it works. We certainly did not make this stuff up ourselves. These are standard tools based on statistical rules that are used in industry.

I know you are convinced you are right about what ever it is you are talking about (since you keep repeating it over and over and over with little explanation). I do believe that you believe you are right. But since you have made no case other than to say we are wrong, and that we should take it on faith simply because you said it, I guess we will continue to trudge along without your blessing. I would hope at some point you would explain your assertions (what formulas you are using or what method you are quoting would be useful). After all, this is not meant to be pop quiz for anyone.

My recommendation (completely unsolicited) would be to come into the fold since your big ol' brain is certainly a resource, and do a little better job of wholly explaining your basis for debate. What we appear to have here is an argument (of Monty Python fame) and not a debate. You are simply disagreeing with little or no explanation. A second recommendation would be to assume that you are not the only qualified voice in the debate, there are some other sharp guys out there.

I hope you take this post with the spirit that it is meant. If not, I am sorry.

godzilla · Oct 18, 2004 06:54 PM

RE: Scoring Method#85 source

>>Absolutely, we should be comparing methods between judges
>>and aligning them. Are you kidding?
>
>Ok aligning the like methods yes. But compering unlike
>methods and saying one is wrong and the other is right.
>This can only be done when a judge uses NON RB criteria,
>such as me and inverted flight. And no I am not kidding.

Once again, this is showi ng a coorelation, and it is simply an example. The assumption would be that coorelation would show consistency of measurement.

Even a coorelation of 1% is a coorelation.

ferocious · Oct 18, 2004 05:14 PM

#78 source
you just came up with a good list of what should go into the judge's guide. Want to volunteer?

dirtydan · Oct 20, 2004 05:56 PM

edited#94 source
As a general comment, much of this reminds me of a period of time when I had a fellow working for me who went to three-day seminars about every two days.

Almost without exception, he would come back from a seminar just bursting with "new" information that he had learned over a two- or three- or four-day period and which he just assumed the rest of us had never been exposed to for some reason.

He began to rather annoy me after a fairly short period of time, and so I took to sending him out to deal with those in the company's employ who really were accomplished experts in various fields. People like Brett.

This was not a totally acceptable move, but just before I had to actually deal with the situation, our Instant Expert moved on. Seemed that we had paid for enough seminars that he regarded himself as being more than we could afford.

And he was right. It was a win-win deal.

Dan

godzilla · Oct 20, 2004 06:30 PM

edited#97 source
zzz

50plusAirYears · Oct 20, 2004 05:53 PM

#93 source
>Keep in mind, this is all fake. It is also just for fun and
>something for me to do to practice using my new software.
>Don't anyone go nuts here.

Sounds like not everybody here read the complete message!

Entry into Six Sigma methodology is hard enough if all you do is take a class and work their examples. It's also a use it or loose it for a lot of people. Unless your job requires Six Sigma applications for a large part of your assignments, you almost need to make up your own problems to keep this toolbox functional. Keep it up, and remember, "Corburundum Non Illigitimi est!" or words to that effect.

Howard Rush · Oct 21, 2004 02:11 AM

#101 source
Hey, you guys, I got a statistics problem for you. What's a continuous form of the Binary distribution? I want to know for a problem I have at work. I can't justify the use of such a function, but it would take some unwanted jumps out of some calculations.

klelmore · Oct 21, 2004 11:13 AM

#103 source
Do you mean the bonomial distribution? If p = prob of success and N = number of trials, then two criteria need to be met:

Np > ~5 AND
N(1-p) > ~5

In that case, the statistic

(Xbar - p)/sqrt(p(1-p)/N)

has an approximately normal sampling distribution with mean 0 and unit variance.

Is this what you're after?

Kim

Howard Rush · Oct 21, 2004 09:17 PM

edited#106 source
Thank you for responding. In my case, pN is usually less than one. What I'm looking for is typically the cumulative probability of number of successes in 700 trials, given a probability of success per trial of .00003. Plotting it, number of successes is on the x axis, and the y axis goes from 0 to 1. I think that's the binomial distribution, but I want something that also has a value for nonintegral number of trials, if such a function exists. My whole approach may well be bogus.

I haven't looked at the stuff in the above thread. It seems to be a bit quarrelsome, so I took the liberty of derailing it for personal, nonmodeling gain, seeing that it's attracting the attention of statistical people.

Pat Mackenzie · Oct 21, 2004 10:05 PM

#107 source
Howard,
my math is very rusty, but this site shows the derivation of the normal distribution (continous) from a binomial one:
http://mathworld.wolfram.com/BinomialDistribution.html
Pat MacKenzie

Howard Rush · Oct 21, 2004 11:05 PM

edited#108 source
Thanks, Pat. That will take me awhile to chew on.

That Poisson distribution looks fishy.

Dave Simons · Oct 22, 2004 12:38 AM

#109 source
Really, it's pretty bass-ic

Howard Rush · Oct 22, 2004 01:21 AM

#110 source
Maybe, but I'm floundering.

Dave Simons · Oct 22, 2004 01:51 AM

#111 source
I guess that you'll just have to mullet over.

klelmore · Oct 22, 2004 10:05 AM

edited#113 source
Ah. I see. No such function exists as far as I know. We occasionally have to face this, so we simply peform a linear interpolation between the two closest points and carry on.

It's not strictly valid, as the value for non-integer trials is undefined, but to do the problem we have to come up with *something,* and it certainly seems to work well enough.

Kim Elmore

Howard Rush · Oct 22, 2004 02:20 PM

#114 source
Thanks. I think that is a good idea. Now I have to implement it such that it takes a noneternity to run in Visual Basic. I'll still try to understand the reference that Pat sent, although the fish jokes are getting pretty crappie.

Dave Simons · Oct 22, 2004 04:37 PM

#115 source
fish jokes? maybe your confidence in the analysis is mis-plaiced.

godzilla · Oct 22, 2004 05:28 PM

#116 source
>Thanks. I think that is a good idea. Now I have to
>implement it such that it takes a noneternity to run in
>Visual Basic. I'll still try to understand the reference
>that Pat sent, although the fish jokes are getting pretty
>crappie.

Keep it up Howard and I will have Phil beat your *bass* with a HALIBUT!!!!