Stuka Stunt Control Line Forum
Archive, 2000–2021 · recovered from the Internet Archive
Forums › Stuka Stunt Main Forum

We ain't alone in Judging

Stuka Stunt Main Forum · 48 of 48 known posts recovered

ferocious · Aug 29, 2004 01:12 PM

#0 source
If nothing else, the Olympics give us a chance to see how other events do judging. They all seem to have the same problems as stunt judging. I caught the finals of the men's platform diving. It was quite intense and very informative.

For one thing, it really seemed to be pretty easy to judge, at least among the top finalists, who were very good. Since they all dived just about perfectly, it boiled down to: how far from the platform(too far is bad), a straight entry into the water, and size of the splash. Us folks watching had no problem agreeing with the judges.

Was a different story in the women's platform. The one US contender got downgraded not entering vertically. She curved her body a bit forward. But a few dives later another diver bent a bit backwards but didn't get any downgrades. They both splashed a lot. You just can't beat the judges.

One of the Chinese men spoofed the judges. He got really fast rotation in his somersaults by tucking his legs wide apart, so he could pull them in closer. It should be a big downgrade, but the judges stand directly to the side and couldn't see it. Sound familiar to anyone who has stretched the loops or vertical eight a leetle large?

Phil C

Brett Buck · Aug 29, 2004 02:08 PM

#1 source
>If nothing else, the Olympics give us a chance to see how
>other events do judging. They all seem to have the same
>problems as stunt judging

And those problems are:_____________________________?
In other words, just more of the same vague fingerpointing. Must be August again.

There may well be problems ( I know of a few minor issues) - but I doubt you know what they are. Any more than you do about diving. Didn't enter vertically? Maybe they did something else better elsewhere that you didn't understand ( i.e. didn't get mentioned on the TV broadcast, when the announcers were rejudging the dives using slo-mo video from all angles that the judges are not allowed to use). That's why they don't get Olymipic diving judges from amongst the international TV audience, that's why the announcers don't get a vote, and why we don't get NATs judges from amonst the random casual spectators, peripheral competitors, and from other events

And, it's not a "problem" that it is subjective, any more than it's a problem in diving. It's the nature of the event.

I would also note that many of the problems in gymnastics were the result of errors in *perfectly objective* assessments of start values, and also of how many "holds" there were. Supposed to be 3 holds, there were 4, no subjectivity involved. manadory .2 point deduction - that wasn't taken.

You want to help. Here's what you do:

1) Identify a specific problem that exists today. Not examples of people who "should have won", nor apocyphal stories of "they did this at the Podunk invitational in 44BC".
2) Discuss problem with people involved in event, get agreement among event experts that it's a problem (and not the result of misunderstanding that results from being, essentially, a non-participant in the event).
3) Identify a specific solution
4) Assess solution for practicality, logistics, conformance with basic concept of the event. provide analysis as necessary
5) Discuss solution with event experts (i.e. national experts, PAMPA Competition Comittee), get agreement that solution makes sense and actually solves specific problem identified in 1), and doesn't cause other, worse problems, either inadvertently or on purpose
6) Submit rules proposal, with analysis, to CL Aerobatic CB

Most of what I see skips step 1 - and without step 1, there's no chance of steps 2-6 happening. Particularly Step 6.

Brett

Lou_Crane · Aug 29, 2004 02:21 PM

#2 source
Brett,

I hope a great many of us have, or will, read your lengthy comments in the Jul/Aug 04 SN.

You often speak directly (sometimes it sounds bluntly), clearly and on the best solid info you have available -- Your SN piece shows your purpose is NOT to enter personal combat. It also does that better than there's room for in a forum like this.

KUTGW!

\BEST\LOU

Bob Reeves · Aug 29, 2004 08:06 PM

#3 source
First off I completely agree with Brett, if you can identify a problem.

I could be out in left field but I think most that post new ideas are not even thinking about a problem. I believe most are just throwing out ideas, sorta like a corporate brain storming session. Does corporate America still do that? Anyway, all this stuff on judging are just ideas thrown in the ring to see if one of them might just possibly make Stunt better or judging easier. Why is everyone taking this stuff so seriously?

ferocious · Aug 30, 2004 09:12 AM

#4 source
I have identified the problem numerous times in the past. The variability in scores is relatively large so the resulting placings are often likely just random rankings. Add in the facts of ballooning scores, biases, and(at least at local contests) the fact that some judges more or less score by the class and the results are even less reliable.

It's a good system. Almost everybody gets to feel good about the results, at least once in awhile.

Phil C

grzly23 · Aug 30, 2004 11:02 AM

#5 source
Phil,

You and I both have been at the Brodak several times and I have yet to see or experience that problem or problems you allege. And I agree with Bret. The inflation seems to be endemic to the FAI not the US system of scoring. As far as class, I would hope the higher one moves in class; Beginner, Intermediate, Advanced and Expert, your score would get better. Only seems logical.

Tom McClain

Brett Buck · Aug 31, 2004 12:08 AM

#10 source
> As far as
>class, I would hope the higher one moves in class; Beginner,
>Intermediate, Advanced and Expert, your score would get
>better. Only seems logical.

Phil actually has a good point in this regard. It's been a notion, triggered by the habitual and mostly ritualistic practice of throwing out high and low, that you need to try to target your maneuver scores in certain ranges based on who is flying. Otherwise your score gets tossed out and "doesn't count".

Couple problems with that, of course - scoring high and getting tossed out *does* still effect the score (see other thread). But a far worse problem is that it involves pre-judging based on who the pilot is . This makes it much more likely that the "right people" will win - i.e. the people who are expected beforehand to win. Making it far more difficult for an unknown to get a high score he deserves, and makes it far more easy for a name to get thehigh score he doesn't deserve. I make the point in the other post that I would have to say it probably doesn't often effect he results, but if you happen to be the newly-minted expert who flew better than your local universal stunt hero but finished 6th, that seems like a hollow argument.

Brett

Brett Buck · Aug 30, 2004 11:55 PM

#9 source
LAST EDITED ON Aug-30-04 AT 11:56 PM (CDT)
 
>I have identified the problem numerous times in the past.
>The variability in scores is relatively large so the
>resulting placings are often likely just random rankings.

That's what you meant? - There's a good arguable point.

But it was completely and thoroughly debunked, over on Flitelines, mostly by me but also by others. But to repeat my side of the argument:

The basic premise (that the scores are subject to random errors, and that averaging them and then taking the standard deviation suggests they not stastically signifcant) is just not right. There is no "right" score in an absolute sense - therefore attempting to retreive the "right" score through statistical means by comparing the absolute scores is not valid. Nor is there any idea the the variations from the presumed "true" score somehow reflect any recognized standard "noise" distribution.

The search for the "right" absolute score could potentially be solved by first normalizing the score over a very large number of flights - at least then you are comparing apples to apples. But there's still no way to prove, nor has anyone attempted to prove, that the variation follows the definition of random in any distribution. Without that, the rest of the "analysis" completely falls apart.


>Add in the facts of ballooning scores, biases, and(at least
>at local contests) the fact that some judges more or less
>score by the class and the results are even less reliable.

Ballooning - we know how to solve that to a reasonable degree. Compare NATs and W/C.

Biases - "bias" in the sense that someone might cheat for whatever reason - is completely indistinguishable from "preference" for a particular set or type of errors. Evaluation of a set of errors and assigning them deduction is precisely what we get judges for - within reason, there is no "right" answer, and therefore searching for maliciously applied bias instead of an intentionally applied bias (i.e. preference for a better flight) is a hopeless task, at least in ex-post-facto processing.

Example - not real, but hypothetical with real names. Peabody is judging. Windy flies, gets high score on circle. Peabody it highest on Windy, others have him , say, third.

Is it because of Peabody cheating for someone who is nominally his buddy -and therefore maliciously biased?

Or is it because they both live in the Northeast, have a similar background in the event, and they both have similar views of what errors are severe downgrades, and which are minor - which is exactly what we have judges for, and is an intentional bias in favor of what Rich considers a superior pattern? It's not cheating, it's precisely why we asked him to judge in the first place.


Even more to the point, please explain how you can distinguish between the two cases after the fact? Tossing him out "high" is both tossing good information, and doesn't really work anyway (see below)

Score-by-class. Bad practice, shouldn't happen, if it does it should be trained out, or practioners not used as judges.

I would, however, wager it only makes a difference in very rare cases (where someone comes out of nowhere, or is wildly mis-classified). I think we have adequate information about the NATs to eliminate this as a serious problem - most advanced pilots did indeed finish below the Open, but I know for almost certain that many of the judges had no idea whatsoever what class most of the people they judge flew - far too many pilots to keep track of and I doubt anyone cared, anyway. It does matter if you think you can use the absolute scores to determine advancement, and I think that idea has pretty well been debunked.

Of course, the easiest solution of all is to get rid of skill classes at major contests. J/S/O, sounds fine to me, bring it on. At least for 2 more years until Rob Gruber gets old enough - then I'll have to retire, anyway.

I might also add that this is an offshoot of something that I actually do admit is a less-than-deal feature used in many big contests - throwing out high and low. It's a move originally intended to eliminate cheater high scores from effecting the results. Of course, it didn't take long to realize that it really didn't eliminate cheater high scores from the results. It could only possibly work if there was a correct absolute value for a flight, which there isn't. Once you give up on that notion, it's obvious that given a range of "non-cheater" scores, having one guy go out high definitely effect the results, by leaving the next-highest legitimate score in.

That's exactly what happened in China - the Chinese judge said he couldn't judge his countrymen fairly, so he gave them all 10's so he would go out high and "not effect the results". Of course, this raised their scores because the legit high score that would have gotten tossed otherwise got to stay in. In other words, at best throwing out high and low would only reduce somewhat the effect, not eliminate it.

You can't eliminate cheater scores or the effect on the results by throwing out high and low - period.

Everybody in any position of responsibility understands that this is the case. Problem is, how do you stop doing it? Once you stop, the general consensus is that there would be all sorts of complaints that we are now beginning to allow cheating - since we got rid of the "solution" . Analysis of NATs and TT scores indicates that the ranking would not frequently be altered by keeping all the scores, or tossing high or low, and then never by much. So it is kept - not to accomplish any useful purpose aside ftom making people feel better that we are "doing something" about cheating.

And before you say it - yes, I have discussed this *quite extensively* with all the right people and there is no signifcant disagreement over the situation. I did not propose a change to use all the scores because just like tossing high and low doesn't really effect anything most of the time, it doesn't effect it much to put it back, either. It does tend to normalize the absolute score, but that's not relevant.


Comments?

Brett

Dave Simons · Aug 31, 2004 12:41 AM

#11 source

".....compare NATS & W/C": You can't, they're different animals.

<sigh>

ferocious · Aug 31, 2004 11:40 AM

#12 source
Sorry Brett, but you haven't "debunked" anything. Your argument that there is no "right" score is meaningless. If there is no "right" score in an absolute sense, then my 250 score beats your 525. If neither score is "right in an absolute sense" then the whole scoring system is meaninless and pointless. We both intuitively know that is not true. The scores do mean something.
Judging, of any sort, is a well-researched area of psychological testing. There are correct answers and proven procedures to find them. We just need adopt the best methods. You see it everyday in political opinion polls, taste-testing a piece of chocolate, ad development on Madison Avenue, judging stunt flights, and test marketing of various products. You ask the "judge" to assign a score to something.It is a fact of life that, given the same stimulus, the judges will give varying scores. So all the statistics of standard deviation, means, etc. apply. It has been shown time and again that in a task like assigning a score to a maneuver judges do vary randomly. There are minor discrepancies in that judges tend not to use the ends of the range(few 10's or 40's) and they tend to cluster scores around divisions, or develop pet numbers, but that doesn't really affect the overall results.

It is another basic fact of life that more judgements will give more reliable, accurate answers. Statistics tables don't even consider scores that are averaged from fewer than 5 judgements. If you work through the numbers, five judges have to track each other exactly(all go up and down the exact same numer of points on different maneuvers) except for one or two maneuvers where they differ by a point, in order for score differences of less than 5 pts. to likely actually be different.

Once you get sufficient number of judges(somewhere over 20 or so) almost all the "problems" in judging go away, due to the law of averages. Small score differences reflect real differences in pilot performance. The effects of any kind of bias between judges tends to disappear in the average. One judge simply can't have a big effect on a 20 item average, unless they are blatantly out of line in their scoring. Judges can be rotated in and out of the lineup without affecting the outcomes.

I've never advocated throwing out any scores, high, low, or mine. Scores are the only data we've got and they all should be used. Unless, of course, there is some obvious problem- the judge forgets to score a maneuver, the judge admittedly scores mistakenly, the judge admits to scoring improperly to bias the results, etc.

You want examples of bias. Here goes. I won't name names because I don't remember them and they really aren't to the point. I put up two flights before one pair of judges. They were 400-ish flights. I got 236. Later on I use the same plane in a different event and put up two 400-ish flights. I get 411. Apparently the first judges did not like to see a "combat" profile in amongst the full stunt ships. The second set of judges judged the maneuvers. With an adequate number of judges(such as the contestants themselves) this kind of bias won't affect the results much, if at all.

Then you have the class problems. Someone mentioned it here in connection with one of the Russian flyers where the judges did not score obvious mistakes. I put in a flight, nail the takeoff and landing, and make some really big mistakes on the loops and the eights. I look at the score sheets and see 31pts, +-1 on every maneuver. Obviously the judges don't know their job. No matter how dedicated the judges are, if they don't have a good description of what exactly they are supposed to be judging, and know how to do it, the results will be useless. If there is no agreement on what kinds of mistakes there are, and how important they are, the results aren't going to be useful. adding more judges won't improve things either. Kudos to the Olympic judges I saw in action last week- they did have a consistent set of criteria and they seemed to apply it very even-handedly.

With all that, I can still go to a stunt contest and have fun. I know where I made the mistakes. I know if I flew better or worse than last time. I enjoy occassionally nailing a maneuver. I like the feel of a really good flight. Sometimes the judges agree and hand me a trophy, but that is really neither here nor there.

Phil C

Brett Buck · Aug 31, 2004 04:35 PM

#18 source
>Sorry Brett, but you haven't "debunked" anything. Your
>argument that there is no "right" score is meaningless. If
>there is no "right" score in an absolute sense, then my 250
>score beats your 525.

Not at all, and the essence of the problem. If you get 250 and I get a 525 on the same circle on the same day from the same people, then I get 1st and you get 2nd. But if you get a 250 at one contest, and 525 at another, it means absolutely nothing at all about the relative quality of flight. By the same token, and more on point, if you get a 250 from one judge and 600 from another judge on the very same flight, it ALSO doesn't mean anything in and of itself. It only matter that the first judge had you ranked relative to the other fliers correctly.

That's why your idea is not mathematically or conceptually valid. If Bob Foster gives me a 575, McClellan gives me a 525 and Mark Overmeier gives me a 475 you can't in any valid way peform statistical analysis saying that the right score is 525+-50, and therefore anyone within that range is statically insignificant. That's complete statistical nonsense, as this calculation presumes that the difference is the result of random variation - which is most certainly does not.

More later...

Brett

ferocious · Sep 02, 2004 11:56 AM

#23 source
LAST EDITED ON Sep-03-04 AT 12:28 PM (CDT)
 
With three judges, each judging 15 maneuvers, you can do an analysis of variance and produce average scores for each judge vs each maneuver vs each pilot and determine a standard deviation and confidence interval for the scores. How well the three judge's scores track each other will give the standard deviation showing the randomness in scoring due to the judges. Three judges is plenty to get good standard deviations. The problem is that with so few judges the statistics show that the standard deviation in the scores will so wide that it takes huge differences to be confident that the scores are different.

This is VERY hard for people to accept, but it is a scientific principle just as much as the F=MA equation that drives the planes. Unlike judging the weight of a couple of rocks or the trajectory of a fly ball, people have very little FEEL for how statistics works. Even when you have trained people who have been working with this stuff for years, they still get it wrong. They see a chart, see the 4.1 and 3.9 and act like that is a real difference and ignore the little error bar that shows a range of +-2.1. Happens all the time. They ignore the fact that the overlap in the scores is so broad there is no difference. Then they go and make a marketing decision and lose ten million dollars in wasted advertising.

Using three judges, the current 10-40 scoring, and the usual whole point changes that most judges use, the judges have to track each other PERFECTLY, except for a 1 pt difference on one or two maneuvers, in order to get the variation in the scores small enough that a 5pt score difference means anything.

Phil C

godzilla · Sep 02, 2004 01:58 PM

#24 source

>Using three judges, the current 10-40 scoring, and the usual
>whole point changes that most judges use, the judges have to
>track each other PERFECTLY, exept for a 1 pt difference on
>one or two maneuvers, in order to get the variation in the
>scores small enough that a 5pt score difference means
>anything.

Good point.

The City Smasher

Brett Buck · Sep 02, 2004 02:19 PM

#25 source
>With three judges, each judging 15 maneuvers, you can do an
>analysis of variance and produce average scores for each
>judge vs each maneuver vs each pilot and determine a
>standard deviation and confidence interval for the scores.
>How well the three judge's scores track each other will give
>the standard deviation showing the randomness in scoring due
>to the judges. Three judges is plenty to get good standard
>deviations. The problem is that with so few judges the
>statistics show that the standard deviation in the scores
>will so wide that it takes huge differences to be confident
>that the scores are different.

Nonsense. Of course you can compute various statistical parameters for any data set. That doesn't mean that they are valid and indicative of anything, which they most assuredly *are not* when the data set is doesn't meet the requirements for which statistics apply. In this case the very first assumption in deriving equation in any statistics book is "for N>>1" (sample size much greater than one). Three or five is not >>1, statistical mean derived assuming that therefore do not apply. QED, right there.

But I'll continue with some of the other observations in the first page of elementary statistics books.

>This is VERY hard for people to accept, but it is a
>scientific principle just as much as the F=MA equation that
>drives the planes.


Statistics is indeed a perfectly valid mathematical field. But you are not using them correctly, as the underlying precepts of statistics do not apply to your example. For instance, comparing items that contain differences not arising from random, statistically recognized, "errors". I use the word "error" advisedly, as it's not really error, but difference.

Say you want to measure the length of a pencil. You measure it dial calipers (the same dial calipers) 100 times. You readings ranging from 5.98 to 6.02". The average is 6.0001, and the variance is .001. That's a statistical measure. So you would conclude from this that the pencil is somewhere between 6.0001" +-.095, 3-sigma. Meaning that there is a ~99.7% likelihood that it's *really* somwhere between 5.9505 and 6.0949.

Except, oops, you forgot to zero the dial first, unbeknownst to you, and in fact all the readings are biased by .2". So you have a very precise, but grossly inaccurate, measurement. Maybe you notice that an fix it. But you bought the plastic calipers from the local bargain tools bin, and the gear pitch is not really right - say, the nearest metric equivalent of the right gear pitch because that's what they had when they designed it - and all the measurements, even after correcting the erroneous offset, are subject to a *Scale factor* error. So you have an even more grossly inaccurate measurement to very high precision.

But suppose you want to, instead of finding out how long in inches a particular pencil might be, you want to find out which of several pencils is the longest. Your bias and scale factor errors are not relevant anymore. You can repeat the measurement until the precision gets small enough to be smaller than the mean difference between the closest two averages, and then you have a high confidence which is the longest. You are still grossly inaccurate about the absolute length, but which is the longest is proven to any desired degree.

Then someone else comes up with a digital caliper measuring inches, recently calibrated against a standard, and another dial caliper marked in metric. also recently calibrated.

You will get your precise but erroneous absolute answer of 6.0001, but your determination of which is longest is completely accurate.

The calibrated digital caliper gives you a equally precise determination of which is longest - probably no better or worse than yours - but the absolute length measurement is a much more accurate 5.75.

The metric caliper will also provide an equally precise measurement fo which is longest, but an absolute length reading of 146.05 millimeters.

All are equally accurate in determining the longest pencil, with statistical significance.

I would think that the analogy with stunt judging is obvious.

You method is to average a sample of 3 readings and take the variation or standard deviation of these three readings that are not only not the result of random variation, not only with an sample size violating the rules for which standard deviation is derived (N=3, N NOT>>) but also in different measurement basis.

It would be like taking the pencil readings, averageing them all together (6.0001+5.75000+146.05)/3 to get a average length of 52.6000+-80,9, 1-sigma, and therefore concluding the pencils were all the same length (and incidentally in this case, allow the possibility that some of the pencils don't actually even exist - since there's a ~67% chance that they have negative length). This is not valid statistical analysis, it's gibberish.

Same applied to stunt. Judges *absolute scores* for flights, and in fact entire career averages, *are not all the same* from one to the other. It's like comparing metric and inche measurements directly. And we are not really asking them to all come up with, say a 525 on a particular flight, and deviations from that value are *not necessarily errors*. All we should be asking them is to rank the fliers (or find the longest pencil).

I wish I had cut-and-pasted the other reply from Flitelines, since you argument is identically invalid to the last time - and thus the response it still applicable.

>Unlike judging the weight of a couple of
>rocks or the trajectory of a fly ball, people have very
>little FEEL for how statistics works. Even when you have
>trained people who have been working with this stuff for
>years, they still get it wrong. They see a chart, see the
>4.1 and 3.9 and act like that is a real difference and
>ignore the little error bar that shows a range of +-2.1.
>Happens all the time. They ignore the fact that the overlap
>in the scores is so broad there is no difference. Then
>they go and make a marketing decision and lose ten million
>dollars in wasted advertising.

Right, and if they didn't properly understand how to properly use statistics, they would do better guessing.

I am not arguing that statistics as a mathematical concept is invalid - in fact it is valid and I use it every single day (on things that are a lot more costly than tens of millions of dollars - by far- and approve analyses using *far more sophisticated* methods than you are using here although that's entirely irrelevant to this discussion). I am arguing that you are applying them in a grossly invalid way and thus drawing a wrong conclusion.
>
>Using three judges, the current 10-40 scoring, and the usual
>whole point changes that most judges use, the judges have to
>track each other PERFECTLY, exept for a 1 pt difference on
>one or two maneuvers, in order to get the variation in the
>scores small enough that a 5pt score difference means
>anything.

See above.


Brett

Howard Rush · Sep 02, 2004 02:48 PM

#27 source
I use statistics at work, too, and I use them wrong.

godzilla · Sep 02, 2004 10:18 PM

#31 source
> You method is to average a sample of 3 readings and take
>the variation or standard deviation of these three readings
>that are not only not the result of random variation, not
>only with an sample size violating the rules for which
>standard deviation is derived (N=3, N NOT>>) but also in
>different measurement basis.

I think I brought this up in the pilot judging thread. The only way to achieve any statistical significance is to have a larger data set. I agree 3 is not a large enough data set.

It is not important to create an *absolute measure* as you described with the pencil. The absolute does not matter, only a reasonable mean. To average even 4 or 5 scores, or worse yet, throwing out two, has not statistical value either, but this the current practice is it not?

The City Smasher

Brett Buck · Sep 02, 2004 10:50 PM

#34 source
>> You method is to average a sample of 3 readings and take
>>the variation or standard deviation of these three readings
>>that are not only not the result of random variation, not
>>only with an sample size violating the rules for which
>>standard deviation is derived (N=3, N NOT>>) but also in
>>different measurement basis.
>
>I think I brought this up in the pilot judging thread. The
>only way to achieve any statistical significance is to have
>a larger data set. I agree 3 is not a large enough data
>set.
>
>It is not important to create an *absolute measure* as you
>described with the pencil. The absolute does not matter,
>only a reasonable mean.

Actually, only a reasonable *ranking* from each judge. That's what we are shooting for.


> To average even 4 or 5 scores, or
>worse yet, throwing out two, has not statistical value
>either, but this the current practice is it not?

True, since this is not a sampling process - one thing that is probably obvious but I didn't mention this time is that instead of sampling a large population, here we are taking the *entire population* - all scores that exist, an using them *all*. Thats not a statistical mean, it's *the exact answer*. Yet another example of how taking the standard deviation is not valid.

Averaging them will get the right rank in most cases I have tried. If the judges had the same "scale factor" then I think averaging is identical to ranking (but I haven't done the math). Actually, almost any way I have seen to post-process the scores (aside from patently erroneous methods like Phil's, or intentionally wrong/facetious methods like handicaps) results in the same ranking in the vast, vast majority of cases. This includes the variety of ordinal scoring where you actually do exactly like the pencil example - rank the fliers judge-by-judge (based on the individual judge's score) and assign the lowest sum of ranks. Not the form of ordinal scoring Kim Doherty proposed!

I agree that throwing out high and low makes no sense - particularly when you don't first normalize the scores and the same "high judge" goes out every time. But once again, I have taken real scores and done it with or without, and it hasn't made any difference even when it's close. It certainly could, but hasn't in any real data set I have tried.

As I mentioned in the other thread. most people involved realize that it really doesn't do anything useful - but how are you going to get rid of it? It's perceived (incorrectly) as a hedge against the wild and rampant cheating that "everybody knows" is going on- and now you tell me that we aren't going to try to find the cheaters anymore? I can hear the squealing and rending of garments even now.

Brett

godzilla · Sep 03, 2004 12:41 PM

#37 source

> As I mentioned in the other thread. most people involved
>realize that it really doesn't do anything useful - but how
>are you going to get rid of it? It's perceived (incorrectly)
>as a hedge against the wild and rampant cheating that
>"everybody knows" is going on- and now you tell me that we
>aren't going to try to find the cheaters anymore? I can
>hear the squealing and rending of garments even now.
>
> Brett

Maybe it is not complaints of cheating, maybe its just poor "tool capability" which leads to inconsistent judging.

Maybe it is not the poor guy's fault either, maybe he has been sitting in the sun longer than he should and is simply not capable of giving 110% judging each flier equally. We all have our limits. I do believe that the judges would have a tendancy to pay closer attention to a "Master flier" than a regular Joe Schmo from nowhere in a long day of judging. Human nature. Hence the "curve" is blown.

The City Smasher

Brett Buck · Sep 03, 2004 01:12 PM

#39 source
>
>> As I mentioned in the other thread. most people involved
>>realize that it really doesn't do anything useful - but how
>>are you going to get rid of it? It's perceived (incorrectly)
>>as a hedge against the wild and rampant cheating that
>>"everybody knows" is going on- and now you tell me that we
>>aren't going to try to find the cheaters anymore? I can
>>hear the squealing and rending of garments even now.
>>
>> Brett
>
>Maybe it is not complaints of cheating, maybe its just poor
>"tool capability" which leads to inconsistent judging.


OF course it's not cheating, there is no actual cheating going on. It's the perception that there is. It's taken as axiomatic, particular in the marginal competitor ranks. But you *cannot* detect it, nor can you eliminate it, in ex-post-facto processing. Any idea to the contrary it a delusion. We ask the judges to offer their preferences - they do so. But there is nothing, repeat, nothing you can do to after the fact that can determine what the motivation for that preference might be (cheating for someone, or genuinely thinking the stunts were better). Repeat, nothing.

So it's not at all surprising that, as I noted (and I have done this several times, as have others), you get the same answer almost no matter how you do the processing. Throwing out high/low or not, ordinal scoring or not, normalizing, then tossing out high/low, etc - it almost always comes out the same.

Phil's theory is of course absurd as applies valid concepts in invalid ways.

>Maybe it is not the poor guy's fault either, maybe he has
>been sitting in the sun longer than he should and is simply
>not capable of giving 110% judging each flier equally. We
>all have our limits. I do believe that the judges would
>have a tendancy to pay closer attention to a "Master flier"
>than a regular Joe Schmo from nowhere in a long day of
>judging. Human nature. Hence the "curve" is blown.

Which is why you have a qualifying sequence, where the fliers who consistently score better fly against each other in progressively smaller rounds, and thus are not nearly as subject to fatigue and "halo". For instance, who do you think has a halo when you have 5 guys, a four of whom are former repeat national champions and/or W/C winners? Let's see, is it Paul Walker, 9 NATs and a W/C - or Bill Werewage, multiple W/C including last weekend - or David Fitzgerald, 4 time and defending NATs champ - or Bob Hunt, W/C and NATs winner, and AMA and PAMP HOF member? And then have some poor schlub with a spotty and inconsistent record finish right in the middle? That's how much the halo counts for, if you do it right.

Inconsistent and halo-induced judging is almost non-existent in NATs and other large contests. Not 100%, but so close there is only occasionally a small anomaly. That a lot of people get frustrated and think there is has nothing to do with the reality. It's a function of aspiring fliers who hit rough patches, far more than it's something really happening in the event. Seen it a thousand times, in fact, and one of those used to be me.

That's why I am so vociferous in trying to counteract the almost implicit idea that the event is somehow bogus - this is BY FAR the biggest threat to the event; a far bigger threat than BOM or not, sport fliers versus competition fliers, PAMPA politics, etc.

I would not disagree that there may be an issue with halo, etc, at local contests with inexperienced and/or minimally trained judges and a very wide range of skills in a round. But I also contend that there is nothing useful to be done about it aside from better training. Certainly, just awarding duplicate 1st place trophies to everyone withing "1 sigma" is utter nonsense.

Brett

godzilla · Sep 03, 2004 09:45 PM

#41 source
For instance, who do you
>think has a halo when you have 5 guys, a four of whom are
>former repeat national champions and/or W/C winners? Let's
>see, is it Paul Walker, 9 NATs and a W/C - or Bill Werewage,
>multiple W/C including last weekend - or David Fitzgerald, 4
>time and defending NATs champ - or Bob Hunt, W/C and NATs
>winner, and AMA and PAMP HOF member?

Not me if I am the 6th...

...but I will still try to beat them. So will everyone else. I don't get your point. Are you comparing the strength of halos, or the halos vs. the masses? The latter has validity.

Once again, no matter how absurd you may think the idea is, I still say the ONLY true solution that ends all the controversy of which you so accurately described is *all pilots judge*. You have nearly made the case for this argument yourself.

Plain and simply these perceptions will never go away no matter how much you might want them too, or shout to the contrary.

The winners will want the system to stay the same, the losers will want it to change. This is normal in competition, and healthy.

BTW I certainly see your points I am not trying to be argumentative.

The City Smasher

Brett Buck · Sep 03, 2004 10:22 PM

#44 source
> For instance, who do you
>>think has a halo when you have 5 guys, a four of whom are
>>former repeat national champions and/or W/C winners? Let's
>>see, is it Paul Walker, 9 NATs and a W/C - or Bill Werewage,
>>multiple W/C including last weekend - or David Fitzgerald, 4
>>time and defending NATs champ - or Bob Hunt, W/C and NATs
>>winner, and AMA and PAMP HOF member?
>
>Not me if I am the 6th...
>
>...but I will still try to beat them. So will everyone
>else. I don't get your point. Are you comparing the
>strength of halos, or the halos vs. the masses? The latter
>has validity.
>
>Once again, no matter how absurd you may think the idea is,
>I still say the ONLY true solution that ends all the
>controversy of which you so accurately described is *all
>pilots judge*. You have nearly made the case for this
>argument yourself.

No I didn't. If "all pilots" meant all competitive pilots who really, truly know the event, then I could probably live with it. If "all pilots" meant everybody who shows up regardless of experience, well - no. They are the worst possible judges. You have to know what a pattern is supposed to look like, you have to know enough to separate guys with far superior skills - and budding fliers and moderately experienced low advanced and expert fliers are probably among the most likely to choose a hero and find them faultless. It's the old "If Mr. Stunt Hero did it, it must by definition be right" effect.

I assume that you took the N>>1 argument as a point in favor of more, or as many as possible, judges. Well, read further down - the variation is not necessarily, or even likely to be, the result of random measurement errors. This is an even more fundamental fatal flaw with Phil's analysis. Judges scores are not intended nor expected to be the same for a given flight, and effectively, you are attempting to apply statistics to measurements in "different units" in the raw.

>Plain and simply these perceptions will never go away no
>matter how much you might want them too, or shout to the
>contrary.

Possibly true. But as a service to the event, I have been and will continue to attempt to debunk the idea. If nothing else, if it keeps the people most prone to beleiving this (rank beginners, frustrated Advanced and low expert fliers) from going off on the tangent of worrying about this, as opposed to pursuing something useful like learning how to fly, how to trim, and how to set up engines. Even if the "halo" theory was entirely true (and I assert that it is not), there's nothing a particular flier can do about it except to fly better.


Of course, if there was a legitimate way to detect it and eliminate it, it *should* be used. But I have yet to see any way to determine malicious intent from a raw number.

>The winners will want the system to stay the same, the
>losers will want it to change. This is normal in
>competition, and healthy.

Blaming one's lack of success on others or "the system" is, to me, highly dysfunctional. Certainly, it doesn't help them succeed.

Brett

godzilla · Sep 03, 2004 09:47 PM

#42 source
> Inconsistent and halo-induced judging is almost
>non-existent in NATs and other large contests. Not 100%, but
>so close there is only occasionally a small anomaly.

How, exactly, would you prove this to me? I am willing to hear an explanation or a proof.

The City Smasher

Brett Buck · Sep 03, 2004 10:00 PM

#43 source
>> Inconsistent and halo-induced judging is almost
>>non-existent in NATs and other large contests. Not 100%, but
>>so close there is only occasionally a small anomaly.
>
>How, exactly, would you prove this to me? I am willing to
>hear an explanation or a proof.

It's an assertion, and one that would be essentially universal among competitive pilots. I assert that these people are in the best position to know.

Brett

ferocious · Sep 03, 2004 01:06 PM

#38 source

> True, since this is not a sampling process - one thing
>that is probably obvious but I didn't mention this time is
>that instead of sampling a large population, here we are
>taking the *entire population* - all scores that exist, an
>using them *all*. Thats not a statistical mean, it's *the
>exact answer*. Yet another example of how taking the
>standard deviation is not valid.

Treating the judges scores as "exact" is also wrong, since we know perfectly well, and honest judges admit, they cannot rank or score exactly the same from maneuver to maneuver or flight to flight. In a contest we are sampling one trio of judges out of the universe of stunt judges, and their scores for one set of loops, out of all the loops that have been scored. You only have to score three flights to realize that "I think I gave Shorty 29. I think Joe's now are a bit better, but do I give a 32 or 33? And is Slim exactly in between, or does he maybe deserve a 32 instead of a 31". Scores are obviously not absolute, perfect measures, with no variability, either as rankings or as some kind of estimate of goodness.

>Actually, almost any way I have seen to
>post-process the scores (aside from patently erroneous
>methods like Phil's, or intentionally wrong/facetious
>methods like handicaps) results in the same ranking in the
>vast, vast majority of cases.

I am not suggesting any "post processing" of the scores. Just asking folks to face up to the limitations of the current system and think about ways to improve it. It is not "erroneous" as one man's opinion says. It is the way many many statisticians handle identical data day to day, trying to get useful information from data that has a lot of variability.

Phil C

Brett Buck · Sep 03, 2004 01:19 PM

#40 source

>Just asking folks to face up to the limitations of the
>current system and think about ways to improve it. It is
>not "erroneous" as one man's opinion says. It is the way
>many many statisticians handle identical data day to day,
>trying to get useful information from data that has a lot of
>variability.

Non-responsive. You have not addressed any of my counter-arguments. And it's not an opinion - it's mathematically unavoidable.

You have not established that variations from a given "perfect" score are the result of random errors following a Gaussian distribution (or in fact any other distribution)

You have not established that a sample size of 3 is much greater than 1, which is the basis for the derivation of variance and standard deviation you are using.

You have not established that the absolute score is the critical parameter, vice rank.

As the last two are axiomatically true, you are simply, objectively, wrong. It's not a matter of opinion.

Brett

ferocious · Sep 03, 2004 12:22 PM

#36 source
LAST EDITED ON Sep-03-04 AT 12:40 PM (CDT)
 
Brett, of course using statistics equations on n=3 is not good practice. That is the whole point. Three judges don't produce enough info to give an accurate result when assigning numbers to a performance. The point in not using less than 5 or so judgements is that the results are almost meaningless. The current scoring system works OK for what it was originally used for- small contest with 10 or so flyers of widely varying abilities. Two or three judges can easily separate flyers who score 540, 495, 432, 388, and crash. The judging method simply can't separate ten flyers of similar ability or give them any meaningful feedback on how they did on the various maneuvers. It's like trying to sort ball bearings with a rock crusher.

We have to come up with ways to improve the results, such as ways to get more judges on each flight, as Brad W pointed out. Or how to better teach judges their job, as Keith R has been trying to do. How to get the maneuver descriptions and judging guide to agree with each other. The changes to the NATS format over the years has eliminated a lot of the judging variablility and luck factor, which is good, although is instead places a heavy premium on consistency from flight to flight.

In your lengthy dissertation on calipers, the preferred method is to collect all the data and analyze the variance by type of caliper, pencil, and tester(some people do a better job than others). The
ANOVA(analysis of variance) procedure assigns the appropriate portion of the variance to each variable. If you're really fussy you can even assign order of measurement of the pencils as a variable and determine if the pencils are getting worn form the measurement process.

In a fairly typical contest- 20 contestants- you have 20*2*15 judgements times say 3 judges. That gives 1800 measurements(maneuver scores), which is way more than a few, and plenty to do valid statistics. Just as a thought excercise, if the judges track each others scores perfectly and don't change the way they score during the course of the day, all of the variations in scores obviously must be due to the pilots and the maneuvers. You can confidently say Joe Blow flew the best loops, or Fred Fritz was better than Joe Blow. Throw in the usual juding variability, and all you could say is that first place at 540 beat the rest who all scored below 500.

Phil C

grzly23 · Sep 02, 2004 03:42 PM

#29 source
This reminds me of the old truism from the engineering world about statisticians. "There are lies, ##### lies, and then there is statistics."

Phil,

Your solution may appeal to you and some others, but it is too complex and detailed. You need to go back to the KISS principle. The way we do it now for PAMPA and AMA may have its flaws, but it is the best that can be done and keep it understandable by all that participate. And best of all, it works. This reminds me of something Sir Winston Churchill said about democracy. "It has been said that democracy is the worst form of government except all the others that have been tried."

Sir Winston Churchill
British politician (1874 - 1965)

We don't need to get into the problems that have just surfaced with the Olympics in their judging. All of the computer statistic systems still did not prevent the wrong gymnast from being awarded the Gold. Complexity was the flaw in their system and its result was tragic. We don't need that reality in our sport.

Tom McClain

Brett Buck · Sep 02, 2004 03:55 PM

#30 source
>This reminds me of the old truism from the engineering world
>about statisticians. "There are lies, ##### lies, and then
>there is statistics."
>
>Phil,
>
>Your solution may appeal to you and some others, but it is
>too complex and detailed.

Plus - it's not valid!

Brett

Ion Muniz · Sep 04, 2004 07:52 AM

#45 source

SNIP

Why didn't he gine 'em all zeros?

Ion

Brett Buck · Sep 04, 2004 11:15 AM

#47 source
>>said he couldn't judge his countrymen fairly, so he gave
>them all 10's so he would go out high and "not effect the
>results".>
>
>SNIP
>
>Why didn't he gine 'em all zeros?


Good point. I wondered about that, too - for about a millisecond!

Brett

Randy Powell · Aug 30, 2004 12:30 PM

#6 source
Well, one point made I think is valid. Stunt is a judged event. So, like diving, gymnastic, ice skaing or any judged event, some of it is open to interpretation. Some elements, as Brett notes, are not subjective. Did you do 3 loops? You did or you didn't. Other elements are less so. How consistent were your bottoms? I mean, no one is standing there with a tape measure so it amounts to the judges ability to, well, judge. Not too hard if the pilot is less skilled. But for top experts, it can be a tough gig. Experience is the only thing to rely on. You see enough patterns as a judge, you come to be able to tell through experience.

Any judged event must ultimately be based on the experience of the judges and their ability to dismiss beliefs and some subjective issues and concentrate on geometry. Some can do that well, others not so well. But they all try to do the best job they can.

Makes me happy.

dirtydan · Aug 30, 2004 01:10 PM

#7 source
Maybe following got missed, maybe not. But I feel strongly about this subject, am seriously on the side of "What's the problem here?!" sort of thinking, have learned to simply reject out of hand many of Phil's cynical comments toward various CL Stunt events. As in sometimes it is actually painful to read input from a friend whom I know to be intelligent and deeply involved in the flying of, promotion of, CL models.

Brett had quite a lot of common sense in his response, mine is a lot simpler, for a change: Just as in *any* problem-solving activity, **you must first make the problem statement**.

Without a specific statement of the problem, something actually written down and easily understood, we're just lashing out at many highly valued, highly competent people. Brett has, many times, defended those working on the Nationals scene. I do so on a local level, the NW U.S. and British Columbia, to wit:

"As noted elsewhere, I have not been keeping real close track of who is saying what in several different threads devoted to judging, judges, new schemes, old schemes wrapped in new bunting, etc. ,etc.

I do not keep close track of these discussions as they seem so pointless. What we've got is working to a real high standard, I am perfectly happy with the system(s) developed over many years and by highly qualified people who were motivated to do the right thing for our event.

These discussions also frequently begin with the premise that something is wrong, therefore one of our most damaging instincts, that of feeling we must "do something" takes over, often leading to disastrous results. Reference: Law of unintended consequences.

Many people I respect tremendously maintain that the system used at U.S. Nationals each year results in the fairest and most accurate outcome possible. Good enough. Many congratulations to all involved.

On a local basis--NW U.S. and British Columbia--the situation is quite similar, at least as to the results from each contest. I cannot remember *any* scoring disputes which went beyond a mumbled word or two during my entire involvement in CL Stunt.

The best, very best line in this regard was Leo Mehl's comment of years ago, "Look, they gave me a senior citizen's discount on my score!" While there was a touch of grumbling there, it was obviously a light-hearted comment.

While anecdotal evidence is sometimes the worst kind, I've got some from the 2004 VGMC Western Canada Stunt Champs. We were through first round of Classic and while Will Reeb was flying well, Keith Varley was as well and Bob Smiley was working on moving up, it was my opinion that my Super Combat Streak, 20FP for power, was putting it on 'em all.

Yet the score board showed me lagging behind all the above competitors.

While it seemed a little odd, I have enjoyed such a long string of having my feelings of each flight so closely matching posted scores there seemed little reason to ask questions. And the second round was coming up.

Chris Cox was one of the judges for Classic and upon completion of the event he too was a little taken aback by the finishing order. But not enough to cause any sort of controversy.

As it turned out, there was some funniness going on in tabulation. Futzy calculator, some confusing short-hand from the judges. Chris went back and tabulated all the scores once again. The resultant score board looked as if small children had been involved, but he got it right, and there were some fairly significant changes in the final order.

The point is that things go so superbly with our judges that even a situation that fairly obviously begged for scrutiny was regarded as a minor thing. I didn't even add my own score sheets, a fairly common tactic prior to any complaints being made, it came down to one of the judges stepping in to correct what he saw as a questionable result.

When one operates for years and years in such an environment, it is so very hard to place any credibility at all in discussions which begin with a problem statement when there is no problem to state, which must assume some conduct which is less than sportsmanlike from a group of people who pride themselves on a high standard of conduct.

My many, many thanks to all the judges who make this event not only possible but meaningful for so many of us."

Dan

godzilla · Aug 30, 2004 03:31 PM

#8 source
>I do not keep close track of these discussions as they seem
>so pointless. What we've got is working to a real high
>standard, I am perfectly happy with the system(s) developed
>over many years and by highly qualified people who were
>motivated to do the right thing for our event.

>Many people I respect tremendously maintain that the system
>used at U.S. Nationals each year results in the fairest and
>most accurate outcome possible. Good enough.

So, your opinion to ignore all the discussions about changes is purely a second hand reaction? Do you have any personal experiences with our Nationals on which to base your opinion, or are you simply echoing sentiments of others?

If this is the case, it would seem like trying to ban Howard Stern even though you have never personally listened to program.

I will say this Dan. The competitors at the Worlds were NOTHING like had been reported by our returning US personel. I would say the reports were somewhat biased toward the US point of view. I did not see one top level competitor from the non US countries protest, say anything derogitory, fly a junky airplane, or act anything like a total professional. They also fly very, very, very, good, regardless of what the entire stunt community has been told for years. Be careful taking views in stunt second hand.

The City Smasher

Serge Krauss · Aug 31, 2004 12:05 PM

#13 source
Brad-

>I will say this Dan. The competitors at the Worlds were NOTHING like had been reported by our
>returning US personel. I would say the reports were somewhat biased toward the US point
>of view.

I'm sorry, but I really disagree. I believe that the second-place Chinese entrant did not fly nearly as well this time as the U.S. team reported in 2002. His team mate flew well and pretty well matched their characterizations. Many of the foreign entrants flew well; one should expect this, and I did - BASED ON what Bill Werwage told me at our field. I still feel that Fancher and Walker were scored way too low in the second round. Based on the crowd reaction from behind the judges, they doubtless expected high scores. As usual, the U.S. team members were polite in their reactions; Ted only said, "I know I can fly better, but not 200 points better." I've never elicited what I consider a biased comment from any team member about results.

Werwage very definitely changed his style to suit what he felt were the judges' expectations/preferences. I have seen him fly what I consider significantly better patterns as a matter of course from our grass circles. He does not as a rule fly as large loops (especially in overheads) or square corner radii as he did this time. In fact, his characteristic square transitions to horizontal are eerily precise and quick. These were missing and I think this to be intentional. His intersections also were not as good as usual - perhaps from having to deviate from his acustomed timing - but NO ONE's intersections were uniformly good at this year's WC's on the 2nd and 3rd round flights I observed from behind the judges (only watched the U.S., Chinese, Remi Berringer's, and a couple others from this vantage point). So, comparatively, THAT was not an issue either.

I also feel that what has been said in the intervening year about interpretation of rules PLUS the effect of finally having an American crowd present to make their own comparisons should have had some effect on flight styles. I was happy to see that some European and asian entrants tried for tight corners. However, it was obvious to me that again tight corners were not rewarded, and before this goes to its previous point, I have to say that Berringer's corner radii were obscured by the overpitching of his fuselage. I tried to observe the actual motion of the c.g., but could not really see a difference from other entrants' corners. It is tough to actually see this, and perhaps Richard's DVD's will be useful in deciding this "issue".

Anyway, I saw nothing that leads me to mistrust what Keith, David, Todd, and Bill said - publicly or in private to me - after the 2002 WC's. They were circumspect, understated, and respectful in their utterances. I trust that what you have said is your own heart-felt belief. However, with all due respect, what you have reported that you saw this year has often conflicted with what I am sure I saw. At any rate, your word "NOTHING" above is an absolute that I think you might want to rethink.

SK

Serge Krauss

godzilla · Aug 31, 2004 01:18 PM

#14 source
>Anyway, I saw nothing that leads me to mistrust what Keith,
>David, Todd, and Bill said - publicly or in private to me -
>after the 2002 WC's. They were circumspect, understated, and
>respectful in their utterances. I trust that what you have
>said is your own heart-felt belief. However, with all due
>respect, what you have reported that you saw this year has
>often conflicted with what I am sure I saw. At any rate,
>your word "NOTHING" above is an absolute that I think you
>might want to rethink.
>
>SK

Whatever Serge... you seems to pick out what you want to hear.

For many years the grapevine was that the US was leaps ahead of the WC entrants in terms of flying "correct" patterns and equipment. I remember one the US fliers saying that the Chinese fliers flew "edged rounds" for squares, it was in the PAMPA mag, I have it in the bathroom, I have read it many times.

You mentioned only a few WC names. The Japanese fliers were incredible, and should have stayed to fly the Nats. The Yatsenko brothers were also extremely impressive in flying and equipment. The Berringers were impressive to say the least, and Remi's Sportster was FRONT ROW QUALITY without a BOM. Remi's father declined to fly his #1 model to fly a TWIN 4 STROKE, obviously taking him out of real consideration.

The event for you appears to be simply about the final placings, the event for me was more wholistic. The foreign WC group in general was stunningly impressive for me and all of my friends.


The City Smasher

Keith Renecle · Aug 31, 2004 03:35 PM

#15 source
Hi All,

Serge has made some very valid points that I would like to elaborate on. The biggest point that always concerns me, is that it is necessary for someone as good as Bill Werwage to have to fly a pattern that is not close to the rule book definitions in order to win! Is this what we all want? I read about the challenge of trying to figure out what the judges are giving good scores for, all the time, in magazines, and indeed on this fine forum. It is believed that this is just part of the stunt scene. “Always has been…..always will be” Now, to some extent, I can agree that there is some sort of weird challenge in there somewhere, but I reckon that at world champs level, this is a sick state of affairs.

Why do we need rules? We may as well print very basic rules with very little detail. How’s this “3 round Inside loops with bottoms somewhere around 5 ft.” or “2 Inside square loops (no detail at all)” etc. etc……This sounds a little “tongue in cheek” , but if the judges are giving good scores for loops that are obviously way over 45 degrees, and square corners that are very soft, and way over the limit, why do we need all the details in the rules? Paul Walker’s corners are always aggressive, sharp and precise. The FAI judges guide states that the sharper corners should be scored higher. Paul did not get what he deserved.

I feel qualified to pass comment on the FAI rules, as I have been involved in South Africa, since the new rules were begun. I felt that instead of always complaining and whining about bad judging, I would try to find some better way of helping to train judges to see the perspective of shapes on spheres. I studied up on 3-D graphics, and have been reasonably successful in developing a 3-D simulation of a stunt pattern that is flown 100% to-the-rules. I soon found out that the rules have many problems in the definitions. The first problem came in the squares. The old rules called for equal length sides, the proposed new rules called for sides that are vertical to the pilot, which makes the tops much shorter. When I got to the clover, I simply cannot step through it as per the rules, and apply it to the sphere. It is geometrically incorrect. I have written to so many people around the world, and I have found that there are as many understandings of the same story. Some said that we have to make the squares look “square to the judges”. Others said that the sides must slope outwards. The old judges guide said that the sides must slope inwards etc. In a nutshell, there is very little common understanding of how the maneuvers should look. There are also discrepancies between the FAI and AMA rules. How is it possible then, to have a competition to decide who the best flier in the world is, with this kind of scenario??

What do folks like myself do to prepare for the world champs? All we have is the rule book. Before I arrived at Muncie for the world champs, I wrote to the organizers, for clarity on what the shapes should look like. I was not trying to be funny at all. I stated in a nice manner that the rules are in a state of flux, and also that in South Africa, we are rather isolated from any international competitions. I also checked first with our South African Model Aircraft Association (SAMAA) as to whether our team was allowed to ask for this information. I was told that this is in order, and that I was fully within my rights to ask for this information. The organizers wrote back, and said that I should read the rules. I wrote back and offered more of the valid reasons for the questions, and once again I was told to “Read the rules!” They added that they would do their best to stick to them. Well, I think that this obviously did not happen, because flying to the rules was not going to get the top scores.


I am not pointing fingers at any particular judges here, but I would like to make an important observation. If you see how much training that goes into the preparation of the top pilots, it’s amazing! Incredible amounts of time are spent on training, coaching, preparation of the models, sorting out engines etc. We have to ask ourselves if we have put the same effort into training our judges internationally to have a common understanding of the rules. Looking at judges guides through the years, we have not given our judges the proper tools to judge us with. The proposed new FAI judges guide does not even have the “blow-by-blow” maneuver descriptions in it. I was told that this is because the new rules with the 2-D drawings were self-explanatory. Now I ask you, if the new rules are written from the pilot’s viewpoint (as stated), and the drawings are done in 2-D, how is this of any use to the judges who have to figure out how the shapes should look from outside the circle?

If we are ever going to improve this bad situation, then it is time for us to stand together and reach some kind of consensus on how the shape of our maneuvers should look. We can then create a proper judges training guide that shows the true perspective from all angles. Notice that I did not say that I am trying to “solve” the problems! I use the word “improve”, because I fully understand, and accept, that while we use human judges, we will always have the problem of personal preferences, judging fatigue, the effect of the weather, etc., etc. I sincerely believe however, that we have the opportunity now to do something positive about really good judges training. I have already volunteered my services to do drawings for a new judges guide.

I apologize in advance about not making my sim available on the internet at this stage. It is a bit rough, and I’m still stuck on how to program the clover. There are some of you out there that have seen it, and I have been pleasantly encouraged by the comments so far. As soon as it is sorted out a bit more, I will make it freely available for download on the net.

Keith Renecle

Keith Renecle

Randy Powell · Aug 31, 2004 03:43 PM

#16 source
'zilla,

I wasn't there, so I can't really comment, but have seen quite a bit of footage of the Wch. The styles of European and Asian fliers were markedly different from each other and what I see of US fliers. This isn't to say they were worse or better, but just different. Now, this is from video and I know that the impression that things have live can be quite different; just watch a baseball game on the tube versus going to a live major league game. Very different experience. But I could clearly see a difference in style and approach. I actually liked th approach I saw for most of the European fliers. I didn't care much for the Chinese or Japanese. That is not to say they were somehow inferior or their geometry was bad at all, just a matter of taste.

You could take most of the expert fliers on this board and line them up to watch the World Champs and you would be hard pressed to find any two that agree on the quality of individual flights. It's just human nature.

godzilla · Aug 31, 2004 10:11 PM

#20 source
. I didn't care much for the Chinese or
>Japanese. That is not to say they were somehow inferior or
>their geometry was bad at all, just a matter of taste.

There was one Japanese kid, the one who flew 4 rounds to blow his flight. He flew so small and tight, with tremendous corners. He flew a pipe and an Impact variant. He was SMOKIN' good at flying the Nats style pattern. No 60 degreee pattern here.

The City Smasher

Brett Buck · Sep 02, 2004 02:42 PM

#26 source
>. I didn't care much for the Chinese or
>>Japanese. That is not to say they were somehow inferior or
>>their geometry was bad at all, just a matter of taste.
>
>There was one Japanese kid, the one who flew 4 rounds to
>blow his flight. He flew so small and tight, with
>tremendous corners. He flew a pipe and an Impact variant.
>He was SMOKIN' good at flying the Nats style pattern. No 60
>degreee pattern here.

No one flew 60+ degree patterns that I saw. This is the sort of exaggerations that are making this debate pretty invalid.

The "kid" I believe you are referring to is actually Mitsuru Yokoyama, who was 4th at the 1999 NATs, and is hardly a kid given that he has about 50% gray hair (which is, I guess, better than 25% of all possible hair...). He was the one flying the Impact with piped PA61 and very small maneuvers. He would indeed have done well at the NATs - as did most if not all of the W/C competitors who stayed over. The only guys who I saw in the W/C, who would have not done as well as at the W/C, where the Chinese fliers, Niu and Han particularly. The fairly significant errors in the low-K maneuvers in FAI would have eaten them up in AMA, I think. Assuming of course they didn't, given that they would have known that, go out and improve their performance in those maneuvers. Anybody who can come in 5th at the W/C and then 7th at the NATs is a top flier, by any standard.

I pretty much agreed with the assessement I had heard over the years from the frequent FAI competitors about the quality of flying and the way the contests go - given my knowledge of the way the people involved characterize things. Well less than half the W/C field was "hamburger" by a universal standard including all fliers of all skill levels. But I could easily see why both the US FAI team members and other observers came to the conclusions they did (as did the very knowledgable and skilled US spectator corp). There were things in some patterns that did not meet the standards that top fliers look for (perhaps erroneously, but usually not) to separate the real competitors from the guys you just don't worry about.

Brett

Howard Rush · Sep 02, 2004 02:54 PM

#28 source
Brett may have gray hair, but not statistically significantly gray.

Brett Buck · Sep 02, 2004 10:32 PM

#32 source
>Brett may have gray hair, but not statistically
>significantly gray.

It may well be gray. I don't know, I haven't seen most of it in a long time.

Brett

godzilla · Sep 02, 2004 10:46 PM

#33 source
LAST EDITED ON Sep-02-04 AT 10:51 PM (CDT)
 
> No one flew 60+ degree patterns that I saw. This is the
>sort of exaggerations that are making this debate pretty
>invalid.

Well, Billy flew at least one by his own admission. His last one to WIN was the best 60 degree pattern I have ever seen.

Before anyone jumps on me, just let me say I shook Billy's hand and told him the same thing....and I meant it (of course he had already said he did it on purpose). It was a beautiful pattern. Yes, beautiful. It reminded me how beautiful and powerful our sport can be when the rulebook is taken with a slight grain of salt.

I personally don't necessarily get these little bitty tiny patterns all the time, but I will fly them if I have too.

... and it is not a debate, Brett. I certainly never meant to debate whether the WC entrants were very good. This was obvious to everyone I spoke to except for...two. I did not debate with those two either, why waste the breath? I know what I saw. I know what my friends saw. I also know what I have been reading in past years, which never coorelated to what I saw this year. Maybe this year was fluke, I don't know....

The City Smasher

Serge Krauss · Aug 31, 2004 04:32 PM

#17 source

>Whatever Serge... you seems to pick out what you want to hear.

One of these days, Brad, you will pleasantly surprise me by NOT talking about the people to whom you respond. You will note that this sentence is only the third time in over two and a half years on this forum that I myself have violated this principle of civil discourse. My apologies for a gentle infingement during the second time above. You do not need to insult me. For the record, I assert that your characterization is not only inappropriate, but inaccurate.

>I remember one the US fliers saying that the Chinese fliers flew "edged rounds" for squares, it
>was in the PAMPA mag, I have it in the bathroom, I have read it many times.

This would not be inaccurate. The lower placed chinese flyer flew rounded corners on squares that were recognizable. The higher scored team mate flew "squares" that were scarcely recognizable as squares, having very little length to any "straight" sections. As I said last month, I was disappointed to see this - not happy, but disappointed. I saw a lot more wrong with his patterns; I have stated these in private e-mails (where I'd prefer to keep them) to those who inquired privately. I think that what the team reported from the 2002 WC's was better than what I saw this year.

>You mentioned only a few WC names. The Japanese fliers were incredible, and should have
>stayed to fly the Nats. The Yatsenko brothers were also extremely impressive in flying and
>equipment. The Berringers were impressive to say the least, and Remi's Sportster was FRONT
>ROW QUALITY without a BOM. Remi's father declined to fly his #1 model to fly a TWIN 4 STROKE,
>obviously taking him out of real consideration.

I fail to see this as relevant to anything I said. What's the point?

>The event for you appears to be simply about the final placings, the event for me was more
>wholistic. The foreign WC group in general was stunningly impressive for me and all of my friends.

This characterization is just silly. What I wrote was in response to your characterizations and generalizations. It bears no relationship to my enjoyment of the event, nor to what I see in it. However, next time I want to know what I think about something, perhaps I'll consider consulting you first, especially before I have the temerity to state my own views. Thanks.

Now, I'm wondering whether I'll have to deal with personal characterizations any time I again disagree...

SK

Serge Krauss

godzilla · Aug 31, 2004 10:05 PM

#19 source
LAST EDITED ON Aug-31-04 AT 10:09 PM (CDT)
 
>
>>Whatever Serge... you seems to pick out what you want to hear.
>
>One of these days, Brad, you will pleasantly surprise me by
>NOT talking about the people to whom you respond. You will
>note that this sentence is only the third time in over two
>and a half years on this forum that I myself have violated
>this principle of civil discourse. My apologies for a gentle
>infingement during the second time above. You do not need to
>insult me. For the record, I assert that your
>characterization is not only inappropriate, but inaccurate.

How in the world is this an insult?

First off, you continue to talk about the placings of the current WC's, and who flew the best in the contest. All I was saying is that, in general, the fliers from the non US countries were very impressive (I was simply saying to Dan to be careful of taking things second hand).

I saw nothing wrong with the placings as I never bothered to watch every single flight, as apparently you did. You have you opinion, so be it.

What I was saying, about the fliers themselves and their overall performances was no critique of the standings at the end of the event. Shoot, I went and did stuff, I did not armchair judge the contest. All I know is that there were some tremendously good fliers and many were very, very advanced.

Go back and read who brought up this year's WC's. You did, I did not! All I was saying was that I saw a completely different set of fliers than I was led to believe by some of the things I had read from US fliers IN PAST WC's. Not this WC! You mixed the two somehow. So, it leads me to believe that you read what you wanted to read and decided to SET ME STRAIGHT by applying what I said to the results this year.

The two have nothing to do with one another.

LIKE I SAID BEFORE ON A PREVIOUS THREAD, I have no problem with your analysis of the WC's. Fine with me. However, the idea that these guys aren't very, very, good will never fly with me (this has nothing to do with you).

I suggest you get the DVD from Richard Oliver. It has the scores from all the judges on all the manuevers. Feel free to see the scoring and come to your own conclusions.

The City Smasher

dirtydan · Sep 01, 2004 02:11 PM

#21 source
I will say this Dan. The competitors at the Worlds were NOTHING like had been reported by our returning US personel. I would say the reports were somewhat biased toward the US point of view. I did not see one top level competitor from the non US countries protest, say anything derogitory, fly a junky airplane, or act anything like a total professional. They also fly very, very, very, good, regardless of what the entire stunt community has been told for years. Be careful taking views in stunt second hand.

The City Smash


Godzilla,

This will be harsh...

Referencing last sentence of your missive, *of course* I am, have always been, very careful in "taking stunt views second hand."

To the point where my list of highly credible sources has only five names on it. Brett's is one, mentioned only as he has been pro-active on this issue as it concerns judging standards in use at U.S. Nationals.

Your name is not on this list. No offense intended; it's just a very hard list to get on, and five seems to be a good number, so yours would have to displace that of someone else. Not likely to happen...

When it comes to local contests, local judges, the area in which I fly, second-hand information is not required. And you may consider me to be a highly credible source on this topic.

Dan

godzilla · Sep 01, 2004 04:14 PM

#22 source
>I will say this Dan. The competitors at the Worlds were
>NOTHING like had been reported by our returning US personel.
>I would say the reports were somewhat biased toward the US
>point of view. I did not see one top level competitor from
>the non US countries protest, say anything derogitory, fly a
>junky airplane, or act anything like a total professional.
>They also fly very, very, very, good, regardless of what the
>entire stunt community has been told for years. Be careful
>taking views in stunt second hand.
>
>The City Smash
>
>
>Godzilla,
>
>This will be harsh...
>
>Referencing last sentence of your missive, *of course* I am,
>have always been, very careful in "taking stunt views second
>hand."
>
>To the point where my list of highly credible sources has
>only five names on it. Brett's is one, mentioned only as he
>has been pro-active on this issue as it concerns judging
>standards in use at U.S. Nationals.
>
>Your name is not on this list. No offense intended; it's
>just a very hard list to get on, and five seems to be a good
>number, so yours would have to displace that of someone
>else. Not likely to happen...

This does not offend me.

The City Smasher

Broken Wings · Sep 02, 2004 11:16 PM

#35 source
“The List” lets all just hope that Duh Dirt didn’t draw a heart with an arrow through it next to Bretts name.

Ion Muniz · Sep 04, 2004 08:06 AM

#46 source
You ain't no good, Broken Wings!

LMAO

Ion