Class 95: Answering P-Value Good Uses Challenges

Class 95: Answering P-Value Good Uses Challenges

Reminder: The Thursday Class is only for those interested in studying uncertainty. I don’t expect all want to read these posts. Pease don’t feel like you must. Yet, I have nowhere else to put them. Your support makes this Class possible for those who need it. Thank you. MATH ALERT!

There were only two responses to “good uses” of P-values challenge, both of which from people already skeptical, which I answer here. I could not get anybody who maintains P-values have “good uses” to justify, or even state, what these “good uses” are. Nevertheless, the Challenge will remain open in perpetuity, and I’ll answer any claims that come.

Video

https://youtu.be/YhkPNhPMmJg

Links: YouTube * Twitter – X * Rumble * Bitchute * Class Page * Jaynes Book * Uncertainty

HOMEWORK: Find “good uses” of P-values.

Lecture

Here are the responses and answers.

1. Dr Brown’s

This might fit the bill: Quality control.

Suppose a High-Throughput Industrial plant wants to make sure it’s machines haven’t drifted too far

In automated manufacturing (semiconductors or ball-bearings, profital), a factory line continuously samples physical output to detect machine drift.

  • The Setup: The physical parameters of the baseline process (the Null Hypothesis H_0) are known with extreme precision from millions of prior runs (assumed).
  • The Mechanism: An automated sensor measures a sample batch.  If the sample mean deviates far enough from the target, the calculated p-value drops below a strict pre-set threshold (say alpha = 0.001).
  • Why it’s Legitimate:

1.    The sampling distribution under H_0 is genuinely known, not guessed.

2.    The goal is not “scientific discovery,” but a binary operational action: stop the line and recalibrate the machine.

3.    Over 10,000 daily runs, setting alpha = 0.001 guarantees that you will falsely shut down a perfectly working machine no more than 10 times. It is a pure long-run frequency-capping tool.

Fenix Ammunition out of the Detroit area routinely posts videos from their manufacturing processes, which are clearly designed well. Every casing, for instance, is measured with lasers (I presume) and tallies of various dimensions appear on a screen, with distributions plotted (the real-life ones, not models).

My memory may be faulty, but I believe cartridges that exceed tolerances are automatically rejected. This is clearly an efficient method. No probability of any kind is needed. The rule is set: if dimension X is out more than Y, then reject; else accept.

The can’t be done with pills like Profital, which are made in huge batches. I imagine it could work in part for semiconductors, though it would be costly to check each one in the many dimensions (frequency response, resistance, etc.), and impossible in others (exceedance voltages, which destroy the unit). In these cases, a sampling scheme is a must.

There is no real “null”, though, in any p-value sense. There is instead a list of design specifications with tolerances that are created by the needs of the product and the nature of the machines and processes. There are known goals. The machines (vats or whatever) are made according to these criteria. But machines wear and break, hence the need to sample. Sampling we all agree on, when direct measures of each item cannot be done.

Now we need tolerance decisions, which then give us the outlines of a model, if needed. But it’s not always needed. For instance, if the sampled item exceeds a tolerance (one or several), the probability it has exceeded the tolerance is 1!

What is the probability other measures will also exceed the tolerance? What is the probability nearby (future or recent past) items exceed? That depends on the model we use, and whatever data we deeded important, such as the mean of the samples as you indicated. There are an enormous number of possibilities for models here (fixed or in time), but you get the idea.

An easy model is the one you suggest: simply use the real-life distribution of observed results as predictive of the unobserved. Whichever model is picked, it’s easy to see that probability solves all our problems. We calculate:

Pr(Tolerance decision in new samples | Current sample, old data, tolerances, other information),

which gives the probability of whatever tolerance decision we designate, given the current sample value, any old data we used to build the model, knowledge about tolerances, and whatever other information about the manufacturing process which is important. Simple as that. See the video for clarifications.

In fact, this is even simpler. Because in your scenario, we have a sample, we have the past, we have the decision rule. If the sample exceeds the specified criterion, or criteria, then act; else, not. There’s no need for anything but keeping track of the sample measurements. Why convert them to P-values? The measures are already in the form understood by the machinists, say. We know by design when to stop the machine and recalibrate.

We don’t need P-values, and really don’t even need probability. Probability can still be useful, not in individual decisions, but for predicting how many decisions might be made one way or the other. Like, what is the probability the fraction X of product will pass muster, given the model and so on. Or if the samples are “going” on way or another, we can project when exceedances will be triggered.

Or the probability the machine, or machines, or a specific process, malfunctioned.

That samples are bad we know by design. How they are bad might also be understood, since machines tend to break in the same ways. Probability can be used here, too, to gauge which possible faults are responsible for the given sample exceedances. But even if we don’t know how they are bad, we know (again by design) that they are.

And all this is apart from the mystical parts of P-values, which require infinite repetitions and all that frequentist woo. Those criticisms always hold, but hardly anybody ever thinks of them.

There is, as you might know, an entire field called Quality Control, which goes into all these things. I’ve only sketched the barest minimum here, and have not indented this to be a full discussion of QC, which we’ll leave for the future.

2. Frank Harrell’s

Link:

There is only one counter-argument (against p-values not being useful) that works: I was in a hurry and didn’t have time to do a Bayesian analysis. I’m still looking for an application where data extremeness assuming theta=0 is the best way to quantify evidence that theta>0.

That’s written in statistical jargon, so I’ll first translate, though you might catch of a whiff of Brown’s offering even before I do.

Suppose you have two drugs, Profital and Zxcvbn (“Ask your doctor if Zxcvbn is right for you!”), both in service, it is claimed, as cures for the Screaming Willies, a disease measured on a numerical scale, with higher numbers indicating further-goneness.

A probability model of this measure is then offered, almost always ad hoc, and usually the normal distribution. Here there are two possible models, one for Profital and one for Zxcvbn. Normals have two parameters, a central (where the peak is) and a spread (the symmetric amount of variability around the peak). The spread parameter is usually assumed the same for both models, but two different central parameters are assumed, one for each drug.

Equivalently, one grand model is assumed with that same spread parameter, but where the central parameter is for the difference in outcomes between the two drugs. That’s the “theta” Frank mentioned. If theta = 0, then we use the identical model for both drugs. If theta not equal to 0, then (again equivalently) we have two separate models, one for each drug.

Many other situations fall into this paradigm, some much more complex, but this is the gist, and if we understand this and how P-values fit in, we can follow the rest of them. And in fact, we did all this before in Regression, the right and wrong way, so refresh your memory if needed.

To complete the example: observations of the measure are taken for both drugs, and a P-value is calculated (roughly, a function of the difference in averages of the observations). If the P is wee, it is decided “theta = 0” is certainly false, which thus logically, and inescapably, means “theta does not equal to 0”. And then it is concluded, “The drugs are different in their powers to cause cures of the Screaming Willies.”

Frank well understands this is a fallacious argument, for all the reasons we gave in the Challenge, so I won’t repeat any of them here. His offering is not that P-values “work”, but that this procedure is fast, and makes a reasonable first cut. My answer is that it is indeed fast, and that it can sometimes appear to be a reasonable first cut. Let’s discuss both.

Speed: P-values are no faster, not today, than doing a full probability analysis. It’s all one button on a computer these days, so why not instead churn out

Pr(Cure | Profital, E) < Pr(Cure | Zxcvbn, E)?

Or whatever other question you might have (recall ‘E’ stands for the complex proposition of all the other evidence we’re assuming). Simple! In fact, it is so simple and so easy to understand, that I cannot see any remaining barriers for the adoption of full (or logical) probability, except for the lingering (pagan) belief that probability is in things, that it’s a real force or property of Nature. Which, of course, is false, for all the reasons we covered before.

Reasonable: I have given this example many times, but it’s a must here to repeat. Here is a syllogism:

All men are mortal;
Socrates is man;
Therefore, “2 + 2 =4”.

The conclusion given these premises (a phrase which I ought to write in ten foot high letters) is FALSE. Yet, given other premises, premises we all know and love, the conclusion is TRUE.

When P-values are reasonable, it is because of situations like this. The reasoning which gave the conclusions “our posited cause is likely” is FALSE, a FALLACY, every time, just as “2 + 2 = 4” is FALSE every time in that syllogism. But because, at least in olden days, scientists were good at setting up experiments, and in creating new drugs, about which they had all kinds of other information about their behaviors, information which is ignored in P-value calculations, it was often the case that the drugs really were different, and so the conclusion that the causal powers of the drugs were different was TRUE, but given those other premises P-values ignored. Again, see the video for clarification.

Using P-values in those situations was still wrong, a FALLACY. That fact cannot be evaded. But when there was good information about real causes, the fallacy was not that harmful.

But we are now where we are, with too many scientists, too much money in science, and too much pressure to publish. Now every damned correlation of this with that it taken to be causative because of wee Ps. You and I, dear readers, have looked at hundreds (thousands?) of these papers over the years. So you know the real harms wee P thinking brings.

All my work is free and user-supported. Here’s how to help:


Discover more from William M. Briggs

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *