Class 94: Quantifying The Unquantifiable

Class 94: Quantifying The Unquantifiable

Reminder: The Thursday Class is only for those interested in studying uncertainty. I don’t expect all want to read these posts. Pease don’t feel like you must. Yet, I have nowhere else to put them. Your support makes this Class possible for those who need it. Thank you.

We’re still waiting on people to give their “good uses” of P-values, so I cannot yet post a reply. Perhaps next week? Meanwhile, we return to my Uncertainty book with an excerpt on an important topic.

Video

Links: YouTube * Twitter – X * Rumble * Bitchute * Class Page * Jaynes Book * Uncertainty

HOMEWORK: Review.

Lecture

This is an excerpt from Chapter 10 of Uncertainty. I’ve removed reference links and the like.

The devastation to sound argument by quantifying the unquantifiable cannot be quantified. But it is monumental. Every time a survey with arbitrarily quantified answers is presented as a discovery or as “confirmation” of common sense the road to scientism widens. Here is simple proof of this claim.

On a continuous scale of -3.2 to 113^( 1/3), how would rate the keenness of quantifying the unquantifiable? Or how about something more scientific: on a scale of 1 to 8.3, in units of 1/r where r is a prime number, how flummoxed does the previous question make you? Do you think that a “flummoxedness” of 8 is twice as flummoxed as a “flummoxedness” of 4? And is a “flummoxedness” of 4 twice as flummoxed as a “flummoxedness” of 2? Have we captured all there is to know about “flummoxedness” in this scientifically validated instrument?

Groups of questions are often called “instruments” in order to mimic the prestige things like x-ray spectroscopy have. Social” instruments” are questions with ad hoc numerical values attached to them, given first to a group of volunteers (likely college students since the designers are mostly professors), and then given again to a second group. Ad hoc scores are created from the quantified answers; sometimes the number of separate scores calculated from even a limited set of questions can be quite large. If the scores are somewhat similar between the first and second group, the scientist calls his “instrument validated.” It is then released into the wild—though often with a copyright attached so that would-be administrators of the instrument are made to pay for its use.

There is an old Russian proverb which says the mind of another is like a black forest. Yet modern science thinks it can plumbs the depths of any soul if only enough quantified questions are asked. Though researchers say they are careful about question wording, ambiguity in language, particularly about emotion states and highly charged political questions is high. There is always the danger of questions designed to solicit the desired outcome. Consider “Are you in favor of helping the destitute or would you rather they just die and decrease the surplus population?” versus “Are you in favor of a massive tax hike to fund a new inefficient stream of welfare?”

Nobody disputes that there are levels of, say, happiness. One can be amused, pleased, gratified, sated, satisfied, gleeful, ecstatic, serene, gloomy, depressed, sad, grieved, aggrieved, and on and on. Yet it is only hubris that allows a researcher to say, “How happy are you on a scale from 1 to 10?” and think he has well quantified this complex emotion merely because some people checked off a number. But even this might be okay, this crude, blundering quantifying of the unquantifiable—after all, this is the purpose of all those different words we have for happiness—except that the research must go on and submit his answers to classical statistical analysis. Calamitous over-certainty is the result.

Quantifying the unquantifiable got its real start with Rensis Likert’s (then at New York University) original 1932 paper “A technique for the measurement of attitudes”. Studying this work is revealing. In that paper, Likert (p. 9-10) said “A series of verbal propositions dealing with the same general social issue are assumed to be more or less equivalent, or at least to be closely related so as to permit prediction from a knowledge of a subject’s attitude on one issue to the same subject’s attitudes on other aspects of the same issue….In statistical language, a group factor is assumed at the outset.” He was concerned with measuring “pro- or anti-Negro feeling” on several questions which were to be combined into a “Negro scale” (and a separate “Internationalism scale” which I’ll ignore here). Initiating a practice which would become customary (p. 14) “the attitudes tests were given to undergraduates (chiefly male) in nine universities and colleges.” There were 15 questions fed into his “Negro scale”, each question receiving different numerical weights. Some questions were scaled (in order), such as “Yes”, “?” (he meant indeterminate), or “No”; others were multiple choice. Here is question eight, scored 5 points for (a) down to 1 point for (e):

In a community where negroes outnumber the whites, a negro who is insolent to a white man should be:

  • [5 pts] excused or ignored.
  • [4 pts] reprimanded.
  • [3 pts] fined and jailed.
  • [2 pts] not only fined and jailed, but also given corporal punishment (whipping, etc.).
  • [1 pt] lynched.

Is it really the case that because lynching an “insolent negro” is worth only 1 point and that fining and jailing him is worth 3? Excusing insolence is only five times as merciful as lynching? The answers are obviously not equally spaced in their emotional or cultural content; not for us and not for his nine (yes, nine) students in 1932.

Here is question 12: “If the same preparation is required, the negro teacher should receive the same salary as the white”, with answers (values) “Strongly approve” (5), “Approve” (4), “Undecided” (3), “Disapprove” (2), and “Strongly disapprove” (1). Is it the case that strongly disapproving that a “negro” should receive the same pay as a white teacher is numerically equivalent to lynching a man? It must be, because Likert’s “Negro scale” is a simple average of the answers from all questions, including this one. Except for some adjustment because some questions allowed differing numbers of answers, all questions in the scale are weighted equally. Thus lynching and strong disapproval of equal salaries are by definition morally equivalent.

The percent of answers given to question eight, for example, were: (a) 29%, (b) 42%, (c) 26%, (d) 3%, (e) 0%. Perplexingly, Likert said this and similar responses “yielded a distribution resembling a normal distribution.” That is so only if we take “resembling” in the same sense as “a man resembles a snail” because both are animals. He used this approximation of normality to develop what he called “Sigma scoring” (which is not of direct interest to us) which allowed comparing values of answers of questions with differing number of responses. He summed up the answers, which became the scale.

Likert gave his questions to 8 groups of varying sizes (30 to 100); naturally, the distribution of scores did not match place to place. Variability is expected. But neither did the scores from the same places match the scores on a re-test given 30 days later—there was an 0.85 correlation (calculated in the standard way). Of course, it could be that opinions of the respondent’s changed during this time. Or their attitude towards members of a different race might have remain fixed but their understanding of the question wording change. Or it could be that they paid a different level of the attention from that time to this. Or et cetera. This kind of uncertainty in the answers is never, so far as I have been able to discover, accounted for in any analysis which uses scored questions.

“Validity” to Likert (and his followers) is how well the means and standard deviations of the scores match at different locations (or times). But that is obviously circular reasoning. What is real validity? It should be how well the score matches the underlying truth. That would be how well our scale above measured flummoxedness, or how well Likert’s score measured the complexities of attitudes towards “negroes”. But if our best poets and writers can barely plumb these depths, what arrogance it is to suppose some simple quantified questions can!

As before, it’s worse than this. Researchers rarely report the result of one “scale”, but often how different scales or instruments match one another. If there is any kind of correlation, the emotion or psychological state claimed to be (exactly measured) by one scale is said to either cause or be caused by the emotion or psychological state claimed to be (exactly measured) by the second. What usually happens is that the two scales have similarly worded questions. This is never recognized.

For example, the very widely used Center for Epidemiologic Studies Depression (CES-D) scale in its short form asks level of agreement to inter alia the statements “I felt that everything I did was an effort”, “I was bothered by things that don’t usually bother me”, and “I was happy”. And the just-as-common SF-12 (a distillation of the SF-36) Health Survey has the questions inter alia “Have you felt calm and peaceful?”, “Did you have a lot of energy?”, and “Have you felt downhearted and depressed?”.

Both of these “instruments” have scores. The twelve-question SF-12 claims it can measure 10 separate dimensions of health! One of these 10 is “vitality”. Researchers will model (usually with regression) the relationship between the scores from the CES-D and SF-12, and when “significance” is discovered, they will say something like “Depression lowers vitality” or “Vitality lessens depression.” They will then scour the characteristics of the people measured for clues of how to raise vitality and thus lower depression.

The over-certainty of these works is staggering. Has the CES-D really told us all we know about flummoxedness; or, rather, depression? Has the SF-12 measured with crystal precision vitality? Certainly not. Yet it is always assumed the scale encapsulates every important facet of the emotion or psychological state assigned. The objection that the wordings of questions are similar between two or more “instruments” is never made. The scores are always assumed linear. That magic trick is what allows regression to enter and for “sub-group analysis” to flourish. There are hundreds of standard questionnaires in regular use, which has led to a this-correlated-with-that literature, the very existence of which is used as evidence that researchers know what they are doing. Researchers take comfort that others are doing as they, which is all the proof required that all is well. Once somebody gets a stupendously over-certain theory into print, it is license for more such claims (but this time studied on this or that interesting group that was heretofore “neglected”).

Because of the way they are developed, even inside questionnaires, the similar-wording objection can be made. Researchers designing an “instrument” conjure long lists of questions thought to be related to the emotion of interest. These questions are given to a test audience, as it were, and the questions are successively winnowed by ascertaining how close answers to each other question match in the test sample. This closeness is taken as proof that the questions are measuring the stated emotion. But does “Do you like the color blue” really differ from, “About blue, rate how good it makes you feel?”

For a contemporary example, consider this hotly controversial topic. The paper “Psychoticism, Immature Defense Mechanisms and a Fearful Attachment Style are Associated with a Higher Homophobic Attitude” by Ciocca et al. was picked up by the press and announced with the headline “New Study Suggests Connections Between Homophobia And Mental Disorders“. This press article opened, “Homosexuality was long derided as a mental disorder…but a new study suggests that it might be more likely that it’s actually homophobia that is a sign of mental disorder.” The article quoted one of the study authors (E.A. Jannini) as saying, “After discussing for centuries if homosexuality is to be considered a disease, for the first time we demonstrated that the real disease to be cured is homophobia, associated with potentially severe psychopathologies.” Potentially severe psychopathologies? Sounds like the sort of thing that requires treatment, perhaps against the will of patients.

Now this study asked a few hundred Italian students questions from something called the “Homophobia Scale” as “validated by Wright, Adams, and Bernat” and more questions from the Symptom Check List-90-R (SCL-90-R), “one of the most widely used self-report psychometric tests in the area of psychopathologic symptom assessment”, which provides “nine indexes,” one being “psychoticism”. A regression was ran from the quantified answers and a wee p-value “confirmed” (notice the fallacious identification of cause) “homophobia” described a “significant predictive value of psychoticism.”

Italy is, of course, largely a Catholic country; indeed, 75% of the respondents identified as Catholic. Faithful Catholics are obliged to hold, and many do hold, natural law views of homosexuality which consider homosexual acts as unnatural, sinful, and harmful or (as the catechism has it) “objectively disordered”. Further, while homosexual acts are condemned in the Bible, Catholics are taught to “love the sinner, hate the sin.” These views are central Church teachings, yet there is, as all know, much variation in what individual Catholics believe. Whether these attitudes are right or wrong is immaterial to the discussion here. What is relevant is that Catholic views toward homosexual acts are in part or in whole matters of religion and philosophy, and therefore “homophobia” must be considered in those contexts. Do the scales relied upon by the researchers account for religion and philosophy? No: they do not even come close.

For proof, here are some of the questions on the 25-item “Homophobia Scale”, scored 1 to 5, “Strongly agree” to “Strongly disagree”: 3. “Homosexuality is acceptable to me”, 8. “Marriage between homosexual individuals is acceptable”, 12. “Homosexuality is immoral”, 16. “Organizations which promote gay rights are necessary”, and 20. “Homosexual behavior should not be against the law.” Now each of these contribute to the score in the obvious way towards being “homophobic”. But each also has an answer which is the opposite of homophobia (in the colloquial sense of, say, “hating gays”) but which is in line with Catholic doctrine. We should therefore expect those who are more religious or philosophical to have greater “homophobia” scores but who are not actually homophobic in the sense used by the study authors, which they define as “irrational fear, hatred, and intolerance of homosexual men and women by heterosexual individuals.”

In the SCL-90-R there are 10 questions related to “psychoticism.” The questions are scored from 0 to 4, expressing agreement “Not at all” to “Extremely”. Some of these, if answered honestly and forthrightly by participants, a condition nobody can know, clearly indicate mental difficulties, to say the least. Two questions indicating what most would consider mental illness: “7. The idea that someone else can control your thoughts” and “16. Hearing voices that other people do not hear”. But there are also questions that would just as obviously be answered by religiously or philosophically minded people that do not indicate illness in the context of this study. These are: “84. Having thoughts about sex that bother you a lot”, “85. The idea that you should be punished for your sins”. Don’t forget that these questions were asked immediately after the questions on homosexuality, so that any ideas participants had in this direction were likely amplified.

It is thus no surprise that scores from these two “scales” should exhibit rough correlation—which is exactly what was found: a wee p-value in a model with a small effect and low proportion of variability explained (as is typical), and where the authors gave no indication of having considered alternate explanations. Yet the lead researcher was able to claim, in public, that “homophobia” is a “disease” associated with “potentially severe psychopathologies.” This is not science. It is advocacy or sloppy thinking. There is no third alternative.

We can now see that in the scale of research which causes over-certainty, the runners-up are those researchers whose work leads to newspaper headlines which begin “Science confirms…”. These confirmations will be on such things as “Dressing well improves peoples’ opinion of you”, “Blisteringly hot days are ‘perceived’ as more comfortable than clement days”, “Daytime has more luminosity than nighttime”, and so on endlessly. These studies come out with distressing regularity, all driven by researchers’ need to publish something, anything, and all of which enhances scientism, the fallacy that they only way to know anything worth knowing is if somebody in a white lab coat has certified it. Incidentally, take a moment to spot the real study among the headlines just given. Have it? It’s the first, from the paper “The Cognitive Consequences of Formal Clothing”.

The damage done to clear thinking by pretending batteries of questions adequately quantify emotional states cannot scarcely be underestimated.

All my work is free and user-supported. Here’s how to help:


Discover more from William M. Briggs

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *