http://www.userfocus.co.uk/articles/datathink.html
(pasted here)
Usability Test Data
People often throw around the terms "objective" and "subjective" when talking about the results of a usability test. These terms are frequently equated with the statistical terms "quantitative" and "qualitative". The analogy is false, and this misunderstanding can have consequences for the interpretations and conclusions of usability tests.
Some definitions
Let's take a closer look at what is meant by quantitative and qualitative data. Definitions of ‘quantitative’ and ‘qualitative’ are uncontroversial and can be found in any standard statistics text book. Witte (1989), for example, presents the distinction concisely, defining quantitative data as follows:
When, among a set of observations, any single observation is a number that represents an amount or a count, then the data are quantitative.
So the body weights reported by a group of students; or a collection of people's IQ scores, or a list of task times in seconds, or Likert scale category responses, or magnitude rating scale responses, are quantitative data. Counts are also quantitative, so data showing size of family, or how many computers you have at home, for example, are quantitative.
Witte defines qualitative data as follows:
When, among a set of observations, any single observation is a word, or a sentence, or a description, or a code that represents a category then the data are qualitative.
So "yes-no" responses, people's ethnic backgrounds, or religions, or attitudes towards the death penalty, or descriptions of events, speculations and stories, are all examples of qualitative data. Note that numerical codes can be assigned to represent qualitative responses (for example, "yes" could be assigned 1 and "no" could be assigned 2). However, these numbers do not transform qualitative data into quantitative data.
Note the emphasis on "a single observation". The fail-safe way to distinguish between quantitative and qualitative data is to focus on the status of a single observation or datum, rather than on an entire set of observations or data. When viewed as a whole, qualitative data can often bear a striking resemblance to quantitative data. 57 "yes" responses vs. 43 "no" responses looks like quantitative data. But it is not. Although these numbers are important (and essential for some statistical procedures) they do not transform the underlying qualitative data into quantitative data.
The case of rating scales
Rating scales present an interesting case because they are used to capture subjective opinions with numbers. The ensuing data are often considered to be qualitative. However, rating scales are not designed to capture opinions per se, but rather they are designed to capture estimations of magnitude. Rating scales do not produce qualitative data. Data from Likert scales and continuous (e.g. 1-10) rating scales are quantitative. These scales assume equal intervals between points. Furthermore they represent an ordering, from less of something to more of something — where that ‘something’ may be "ease of use" or "satisfaction" or some other construct that can be represented in an incremental manner. In short, rating scale data approximate interval data and so lend themselves to analysis by a range of statistical techniques including ANOVAs. Qualitative data do not have these properties, and cannot be ordered along a continuum, or compared in terms of magnitude (although qualitative data can still be analysed statistically).
While quantitative studies are concerned with precise measurements, qualitative studies are concerned with verbal descriptions of people’s experiences, perceptions, opinions, feelings and knowledge. Whereas a quantitative method typically requires some precise measuring instrument, the qualitative method itself IS the measuring instrument. Qualitative data is less about attempting to prove something than it is about attempting to understand something. Quantitative and qualitative data can be, and often are, collected in the same study. If we want to know how much people weigh, we use a weighing machine and record numbers. But if we want to know what their weight means to them we need to ask people questions, hear stories, and understand experiences. (See Patton, 2002), for a comprehensive exposition of qualitative data collection and analysis methods).
Subjective Data and Objective Data
A related distinction (and a frequent source of confusion, especially when used in the context of qualitative and quantitative data) is that of subjective and objective data. The "rule" to note is that subjective data result from an individual's personal opinion or judgement and not from some external measure. Objective data on the other hand are "external to the mind" and concern facts and the precise measurement of things or concepts that actually exist.
For example, when I respond to the survey question "Do you own a computer?" my answer "Yes" represents qualitative data, but my response is not subjective. That I own a computer is an indisputable fact that is not open to subjectivity. So my response here is both qualitative and objective. If I am asked to give my general opinion about the price of computers, then my response "I think they are too expensive" will be both qualitative and subjective. If I am asked to report the chip speed of my computer and I reply "2.0 GHz" then my response is both quantitative and objective. If I respond to the question "How easy is your computer to use on a scale of one through ten?", my answer "seven" is quantitative, but it has resulted from my subjective opinion, so it is both quantitative and subjective.
Quantitative | Qualitative | |
|---|---|---|
Objective | "The chip speed of my computer is 2 GHz" | "Yes, I own a computer" |
Subjective | "On a scale of 1-10, my computer scores 7 in terms of its ease of use" | "I think computers are too expensive" |
Confusion often arises when people vaguely assume that "qualitative" is synonymous with "subjective", and that "quantitative" is synonymous with "objective". As you can see in the above examples, this is not the case. Both quantitative and qualitative data can be subjective or objective.
Usability smoke and mirrors
We could put this all down to troublesome semantics and dismiss the matter as being purely academic, but the reality is that clarity of thought and understanding in this area can be critically important. Misunderstanding and — worse — misuse of these terms can signal a poor grasp of one’s own usability data, and may reduce the impact of the results on product design decisions. It can result in the wrong analyses, or in no analysis at all, being conducted on numerical data.
For example, it is not uncommon for usability practitioners to collect subjective rating scale data, and then fail to apply the appropriate inferential statistical analyses. (This is often because they have mistakenly assumed they are handling qualitative data and they assume that these data cannot be subjected to rigourous analyses). It is also not uncommon for usability practitioners to collect nominal frequency counts and then to make claims and recommendations based solely on unanalysed mean values.
Handling usability data in this casual way can reduce the value of a usability study, leaving an expensively staged test production with a smoke and mirrors ending. Such outcomes are a waste of company money, they cause product managers to make the wrong decisions, and they can lead to costly design and manufacturing blunders. They also reduce people's confidence in what usability can deliver.
The discipline of usability is concerned with prediction. Usability practitioners make predictions about how people will use a web site or product; make predictions about interaction elements that may be problematic; predict the consequences of not fixing usability problems; and, on the basis of carefully designed competitive usability tests, make predictions about which choice of design a sponsor might wisely pursue. Predictions need to go beyond the behaviour and opinions of a test sample. In this respect we care about the opinions and behaviours of our test sample only insofar as they are representative of the target market of interest. But we can have a known degree of confidence in the predictive value of our data only if we have applied appropriate analyses. So failing to conduct statistical analyses on both quantitative and qualitative data collected during a summative usability test is a difficult wicket to defend. Such a stance on data analysis could be justified only if we cared not to generalise our results beyond the specific sample tested. This would be a very rare event, and in this case we would not actually be testing a sample but rather the entire population of target users.
Quantitative data are not better or worse, or more or less valuable, than qualitative data. But objective, fact-based data do have greater predictive value than subjective data. Where possible usability professionals should strive to design studies that collect objective, fact-based data.
-- Philip Hodgson, June 17 2003
References
Patton, M. Q. (2002) Qualitative Research and Evaluation Methods (3rd Edition). Sage Publications.
Witte, R. S. (1989) Statistics (3rd Edition). Holt, Rinehart & Winston, Inc.
_______________________________________________________
Rating Scales and Shared Meaning. W Lopez
Lopez W.A. (1995) Rating scales and shared meaning. Rasch Measurement Transactions, 9(2), p.434.
A rating scale is an aid to disciplined dialogue. Its precisely defined format focuses the conversation between the respondent and the questionnaire on the relevant areas. All respondents are invited to communicate in the shared language of the specified option choices (Low 1988).
Ambiguity and uncertainty, however, remain. First, some respondents may not use the rating scale as it was intended to be used. Choosing socially acceptable responses or falling into a response set defeat the purpose of the questionnaire. Second, respondents can only interpret a rating scale in terms of their own understandings of category labels. Lack of clear, shared category definitions invites ambiguity and idiosyncratic category use. Different interpretations lead to inconsistent use patterns.
Traditional statistical analysis, however, mistreats all rating scale observations as precise and accurate communications. Researchers seldom provide for differences in perspectives among respondents. These differences cannot be overlooked if our objective is the pursuit of useful knowledge and sound decision- making. We must recognize the various ways in which rating scale categories might be used and identify those which enable the maximum extraction of meaning. While this involves choice on the part of the analyst, "selective emphasis, choice, is inevitable whenever reflection occurs" (Dewey 1925). Because there can be no knowledge without choice, it becomes the responsibility of the analyst to develop criteria by which those choices can be made.
"Meanings do not come into being without language and language implies two selves in a conjoint or shared understanding" (Dewey 1925). Some level of ambiguity is unavoidable because language can never be exact. Nevertheless, shared meaning cannot be extracted from individual responses unless analysis can identify a common, cooperative mode of communication among all parties concerned.
Rating scale analysis must take the perspective that while a rating scale offers respondents a common language, a tool for "categorizing, ordering and representing the world" (Halliday 1969), it does not by itself make for meaningful communication. Since "meaning is located neither in the text nor in the reader but in their interaction" (Bloome & Green, 1984), we must include a step concerned with discovering, rather than asserting, meaning as we conduct our statistical analyses. Just as readers "must choose between competing interpretations of text" (Bloome & Green 1984) so must the analyst choose between different interpretations of the rating scale in order to find a coherent, shared representation of what is investigated.
A rating scale, like any other tool, "is defined by how it is used" (Halliday 1969). A focus of our analysis must be how the rating scale is actually used by respondents. We must discover which transformation of the initial rating scale categorization extracts the "maximum amount of useful As shared meaning develops, we establish criteria so that we do not ignore the individual, but rather provide a scoring medium through which the dissenting individual's voice may be heard more clearly. We set the stage so that individuals who do not subscribe to our construction of shared meaning can stand out and be noticed. By establishing an explicit commonality among most respondents, we enable the meaning which stems from an individual's unique interaction with an item or a group of items to emerge. The constructive analysis of rating scale data can promote both general dialogue with the group and specific dialogue with the individual. Bloome D, Green G (1984) Directions in the socio-linguistic theory of reading. In PD Pearson (Ed.), Handbook of Reading Research (pp 395-421). White Plains NY: Longman. Dewey J (1925). Experience and nature. Republished in J.A. Boydston (Ed.) John Dewey: The Later Works, 19925-1953, Vol. 1. 1981. Carbondale IL: Southern Illinois University Press. Halliday M (1969) Relevant models of language. Educational Review, 22, 1-128. Low GD (1988) The semantics of questionnaire rating scales. Evaluation and Research in Education 2(2), 69-70. Wright BD, Linacre JM (1992) Combining and splitting categories. RMT 6:3, 233. __________________________________________ and this is: A Statistical Analysis of Teaching Effectiveness http://mrvar.fdv.uni-lj.si/pub/mz/mz17/pagani.pdf ----------------------------------------- Some Cautions: http://www.cs.umd.edu/~mstark/exp101/traps.html Statistical Traps and Pitfalls Statistcs are easy to misuse accidentally, and can also be misused in deceptive ways. This page gives some (by no means complete) advice on how to avoid common problems * Avoid using ANOVA or t-tests with subjective data Avoid using ANOVA or t-tests with subjective data These tests are based on mathematics that require the dependent variable to be measured on an interval or ratio scale (see about statistics), in other word, that your dependent variables act like the numbers you are used to in "normal" mathematics. The problem with subjective data are that on an ordinal scale with values such as 5. Excellent a given subject may not have an equal difference between 3 and 4 as between 4 and 5 (this is what is guaranteed by an interval scale). Since both the t-statistic and F-statistic are computed as a function of differences between measured values and sample means, having an interval scale is necessary for the math to work. Another problem with subjective data is that different people would have different criteria for assigning a ranking of 3 (good) on a subjective scale. There are techniques that allow you to collect useful subjective data (such as the Delphi technique), but they go way beyond what Experiments 101 will cover. And in any case, ANOVA and t-tests should be avoided when looking at subjective data. The null hypothesis is rejected if the probability that it is true is below the significance level set for the experiment. This cutoff is set very low (say 0.05) to reduce the probability of accepting an alternative hypothesis that is wrong (committing a Type I error). If the probability of the null hypothesis is true is a sniggle above the cutoff (say 0.06) we won't reject the null hypothesis. But would you publish a scientific paper claiming a hypothesis is true if you computed its probability of being true as 0.06? An experiment is ideally designed so that (hypothetically) the independent variable(s) represent factors that cause change in the dependent variable(s). However, statistical inference makes no claims of causality. None at all. All that is being done is a computation of the probability that your null hypothesis is true. In scientific research, causality is established by * Having an explanatory theory of causality. An experimental result consistent with such a theory is a good thing, it's just not proof all by itself Even if you have all the above, your model may not explain everything and need to be refined. For example, the model of an indivisable atom was exceedingly useful for 19th century research into chemical reactions,a nd the development of the periodic table of elements. However, it wasn't sufficient to explain why if you had a lump of uranium it would emit radioactive particles and eventually transmute into lead. This lead to a new theory of the nucleus to explain these phenonema. This doesn't mean the old theory is useless -- the old theory probably explains everything you ever did in high school chemistry and most of what you did in college chemistry, (assuming you continued to take any chemistry, that is). However, scientific models attempt to explain all measured phenomena as accurately as possible. Here are two ways that statistically insignificant results can end up being unimportant: * The sample means of different treatment groups may differ by small amounts that have no real world meaning, although they may pass the t-test These two conditions often can happen with large samples. The more data you collect, the finer the detail you can model and make inferences about. A fundamental element of statistics is that as your sample size increases your variance will decrease This can be seen when analyzing the data, at the point where you divide by the variance to compute your t or F value to test. In a wonderful essay "Mad About Measurement" (see References), Tom DeMarco creates the verb "to Limbaugh" to invent a term for the selective use of data, in other words using data that supports your position and discarding any data that has the nerve to contradict you! I'm sure there are good liberals who do this too, but they aren't as visible (in all senses of the word), so I like this term. DeMarco is certainly implying that this is done deliberately, but it is also all too human to have more confidence in data that supports your world view than data that contradicts it. This is an urge that must be resisted when doing a scientific experiment, if for no other reason than to prevent being embarassed when your experiment can't be replicated, or worse, is disproved! This is meant as reminder that statistics is not dealing with exact results, but with probabilities. A place to be particularly wary is if someone uses regression techniques (not a topic covered in this site) to produce a predictive equation, then uses that equation to make an exact prediction of a value for a dependent variable given the values for the independent variables. ACHTUNG! DANGER! This prediction of the dependent variable itself falls into a confidence interval. If you read any scientific paper that graphs an equation derived using regression, look for lines above and below the graph of the equation, ideally with shading between the two. These two lines will give you a sense of the uncertainty inherent in that predictive equation. If they aren't there, distrust the results unless you can get hold of the raw data and run the regressions yourself!
from Students’ Point of View (using scores from questionaires)
* Don't equate not rejecting the null hypothesis with whether or not it is true
* Statistical significance isn't the same as causality
* Statistical significance isn't the same as importance
* Don't "Limbaugh" that data
* Always look for confidence intervals
4. Very good
3. Good
2. Fair
1. Poor
Not rejecting a null hypothesis doesn't mean it's true
Statistical significance isn't the same as causality
* Careful controls for biases. The more carefully you control for biases, the higher your confidence is that your experiment is really focused on the relationship you want to explare.
* Replication. If there really is a causal relationship there, the experiment should be repeatable. This is a lot easier to do properly in chemistry or physics than it is in psychological experimentation, but it needs to be done in any case.
Statistical significance isn't the same as importance
* An independent variable may explain only a small proportion of the total variation.
Don't "Limbaugh" that data
Always look for confidence intervals