Of course we haven't mentioned Bayesianism or Frequentism yet... you haven't said anything about estimating aspects of the probability space from observed data.
Well, we need an independence assumption. And then we have some rock solid theorems and not just something "naive".
Maybe we have to pull the rabbit out of the hat someplace, and with the axioms our place is an independence assumption.
Indeed, the heartburn when trying to swallow frequentist is that we don't have the crucial assumptions to make an average have valuable properties; the sufficient and standard assumption is independence, and then the laws of large numbers, which we can prove just as theorems from the axioms and an independence assumption, tell us that an average does give us what we want to make the whole subject go, in statistics, etc.
So, net, where we pull the rabbit out of the hat is the independence assumption. If we are willing to assume independence, then we don't have to address the philosophical points, or we move the consideration of the philosophical points to considering the independence assumption. But, we are forced: We know that we want averages to be good estimators and know that in practice they often are; since the frequentist approach does not have enough to let us be clear on the properties of an average, we need more; a way to get what we need is the axioms and an independence assumption.
To me the key to probability, and why it is really not the same as just Lebesgue's measure theory, is the role of independence. We can also add in conditional independence. Net, I don't think that we can get very far in probability or its foundations without something as powerful as independence (or at least uncorrelated, which independence implies). The axioms and independence make a nice package to give us a rock solid version of what we want. So, that's why I go for it.
> Well, we need an independence assumption. And then we have some rock solid theorems and not just something "naive".
We have rock solid theorems that the average converges under some conditions. (As an aside, do you rule out the Cauchy distribution in your axiomatic approach?). We don't have rock solid theorems that the average is of any interest. If we're interested in knowing whether a dam is likely to overflow based on historical water levels, the average is of limited use -- for heights well below the highest historical depth, an average will be something that we can at least calculate, but not precisely. If we want to leave room for error (build a dam above the highest historical waterlevel, and do it conservatively so that we'll be above future maixima for the next x years w/ some probability), we can't even calculate an average.
Now what? That's what I mean by "naive." If we're going to try to attack problems that aren't transparently easy (something like stationary weakly-dependent data with finite moments, enough observations that not only the CLT is a good approximation, but that we don't even need to worry about efficiency, loosely defined) then we need to start thinking about the properties we want our estimation procedure to have.
One set of properties leads to something like frequentist statistics, another set of properties leads to Bayesian statistics. Other sets of properties lead to some of the more idiosyncratic branches of statistics that I don't know much about.
> since the frequentist approach does not have enough to let us be clear on the properties of an average, we need more
This makes no sense. Every branch of statistics puts the estimator on top of a probability space, just usually not explicitly. (Because it would be boring.) The t-test has certain properties if the probability space is in a class with....
> The axioms and independence make a nice package to give us a rock solid version of what we want. So, that's why I go for it.
I want to get decent answers for hard problems and have a measure of their reliability. Both frequentist and Bayesian stats are often good for this, and certainly better than "no stats."
> (As an aside, do you rule out the Cauchy distribution in your axiomatic approach?).
Of course not. Essentially the only assumption about a random variable is measurability, and I essentially gave that.
Each of the limit theorems has assumptions, and in some cases minimal assumptions are difficult to consider, e.g., for the Lindeberg-Feller version of the central limit theorem.
For a random variable with the Cauchy distribution, as I recall, the integral of the positive part is infinity and that of the negative part is infinity so that to define its expectation we would have to subtract one infinity from another which, with the measure theory approach to 'improper' integrals, we are not willing to do. So, that random variable does not have an expectation, and the law of large numbers does not apply to an independent, identically distributed sequence of such random variables. Again,
> We don't have rock solid theorems that the average is of any interest.
Sure, we do: Maybe for real valued random variables X and Y and real number a, we are told that
X = a + Y
and want to find a.
Suppose we have E[X] exists and is finite and for i = 1, 2, ..., we have real value random variables X(i), independent and distributed like X. So, then we also have independent Y(i) distributed like Y. Suppose we are told that E[Y] = 0.
Then we can use the law of large numbers to estimate a.
> If we're interested in knowing whether a dam is likely to overflow based on historical water levels, the average is of limited use -- for heights well below the highest historical depth, an average will be something that we can at least calculate, but not precisely.
If we have random variable X that is 1 when the dam overflows and 0 otherwise, then the probability the dam overflows is just E[X]. If we have independent, identically distributed 'samples' of X, that is, X(i) for i = 1, 2, ..., then we can use the weak law of large numbers to estimate E[X].
The role of "historical water levels" here is not trivial to evaluate; the main reason is that we are not sure we have the values of a sequence of independent, identically distributed random variables.
Generally from experience in applied probability and statistics, we should know well that saying what the probability of some particular water level is in the next 100 years is a challenging problem with no royal road to a solution and not a weakness of the axioms, expectation, or the laws of large numbers.
> This makes no sense. Every branch of statistics puts the estimator on top of a probability space, just usually not explicitly. (Because it would be boring.) The t-test has certain properties if the probability space is in a class with....
With the axioms I gave, independence, etc., the Student's t distribution and its applications work just fine as long practiced without mention of either frequentist or Bayesianism.
The reference I gave, Neveu, uses the axiomatic approach I outlined. So do the other famous texts in 'graduate probability' by Loeve, Breiman, and Chung.
So, let's see: I just got out my copy of Neveu and checked the index. 'Frequentist' is never mentioned, and Bayes is mentioned only on one page and there only as a 'strategy' in the "formalism" of statistical decision theory. Net, beyond the simple Bayes rule
P(A|B) = P(A) P(B|A)/P(B)
which is immediate from the definition
P(A|B) = P(A and B)/P(B)
really, Bayes has next to no role in the whole book.
Net, with the axiomatic foundation I outlined, we don't have to struggle, consider, or even mention either frequentist of Bayesianism. That should be welcome good news.
> Net, with the axiomatic foundation I outlined, we don't have to struggle, consider, or even mention either frequentist of Bayesianism. That should be welcome good news.
People have been building on top of the axiomatic foundation that you outlined for decades. This is old news.
You mention Breiman's book on Probability. He has also a book on Statistics. Maybe there is something to it after all?
> People have been building on top of the axiomatic foundation that you outlined for decades. This is old news.
Yup, essentially since Kolmogorov's 1933 paper.
The axiomatic approach I gave is just what is used in Breiman's Probability, as published by SIAM. It's a super nicely written book. I have never seen his book on statistics.
I can't resist quoting Jaynes (from the closing remarks in Chapter 2 of "Probability Theory: The Logic of Science"):
In 1933, A. N. Kolmogorov presented an approach to probability theory phrased in the language of set theory and measure theory. This language was just then becoming so fashionable that today many mathematical results are named, not for the discoverer, but for the one who first restated them in that language. For example, in the theory of continuous groups the term “Hurwitz invariant integral” disappeared, to be replaced by “Haar measure.” Because of this custom, some modern works—particularly by mathematicians—can give one the impression that probability theory started with Kolmogorov. [...]
However, our system of probability differs conceptually from that of Kolmogorov in that we do not interpret propositions in terms of sets, but we do interpret probability distributions as carriers of incomplete information. Partly as a result, our system has analytical resources not present at all in the Kolmogorov system. This enables us to formulate and solve many problems—particularly the so-called “ill posed” problems and “generalized inverse” problems—that would be considered outside the scope of probability theory according to the Kolmogorov system. These problems are just the ones of greatest interest in current applications.
Nothing about this reply makes me reconsider labeling it "naive frequentism." I think you'll find that it's hard to do even bread-and-butter statistical tasks like deciding how many subjects to include in a randomized trial without moving beyond axiomatic probability.
Sorry, it appears that we just are a long way from communicating clearly. I'll try again, a little:
What I outlined are the axioms of probability as started by
Kolmogorov in 1933 and based on Lebesgue's measure theory (of near 1900). So, Kolmogorov finds something in Lebesgue's theory, that is, a particular measure space that has everything we want for probability theory. In this way Kolmogorov shows that we can regard probability theory as just another measure space in measure theory.
Why do that? To get a different probability theory? Not really: Kolmogorov's start just gives essentially the same probability theory we had in 1932. It is just that in 1932, we had to talk about trials, events, probabilities, and random variables without being able to say, mathematically, what the heck they were. To see the importance here, back to near 1900 with B. Russell, etc. there was an effort to redefine everything in mathematics
starting with just sets. Then everything was constructed starting with just the low level ideas of sets. The serious result was axiomatic set theory, e.g., as in P. Suppes, Axiomatic Set Theory. So, numbers, functions, calculus, lines, planes, spheres, groups, rings, fields, vector spaces. etc. were all defined based on sets. And similarly for Lebesgue's measure theory. Then, after Kolmogorov, probability theory was also
defined based just on sets. Whew!
For what you want to do with probability theory in statistics, etc. as you mentioned, Kolmogorov's axioms should be of zero concern to you except you might feel
a little better knowing that the probability theory
you have been using all along does have a solid
foundation just on sets. So, you can go right along
with applied probability as you have been doing, and
Kolmogorov's axioms will essentially never get involved,
never help you and never hurt you.
For more, via the axioms, we can define trials, events,
probabilities, random variables, conditional probability,
stochastic processes, Markov processes, Gaussian random variables, Chi squared random variables, sufficient statistics, distributions, random vectors,
and on and on that you have already been knowing, loving,
and using.
But, maybe I spoke too soon: For Markov processes,
martingales, sufficient statistics, we very much
want the Radon-Nikodym theorem of measure theory
and use it for random variables, and we do. So,
without the Kolmogorov's axioms, we would be somewhat
stuck-o for Markov processes, etc.
> I think you'll find that it's hard to do even bread-and-butter statistical tasks like deciding how many subjects to include in a randomized trial without moving beyond axiomatic probability.
Not at all. There is no "moving beyond". Instead we
proceed essentially as we might have in 1932. That is,
in such work we rarely or never think about measurable spaces, measure spaces, sigma algebras, measurability, random variables as functions on the set of trials, etc. Again, we just continue to do applied probability and
statistics essentially as we might have in 1932.
Or, the Komogorov axioms are down in the sub, sub basement, and we rarely go there. For the rest of the structure, it is just the same or nearly so.
I've studied probability theory, maybe more than you. I like it. But nothing from probability theory suggests why "size" and "power" might be useful properties in test statistics. Without other ideas like those, doing useful statistics would be pretty tough.
I notice you didn't propose a way to choose the sample size for a randomized trial in your reply. I'd love to see it -- using only probability theory and nothing from the stats literature. ;)
> I notice you didn't propose a way to choose the sample size for a randomized trial in your reply.
We're not communicating well.
To respond to your question about sample size, I'd have to look into some of the details of your question. And I
can say now, that question, those details, and any answer
all have essentially nothing to do with the axiomatic foundation
of probability I gave; the answer is the same or essentially
so independent of those foundations down in the deep sub basement of the subject.
Whether something is "only probability" or also "stats"
can be important in practice -- e.g., will find
relatively little about a lot of important
work in statistics in Neveu's book on probability.
And there is a lot of practical knowledge in
applied statistics, e.g., how well
principle components analysis tends to
work in practice (quite well). E.g.,
there is a lot in survey and sampling techniques.
And a lot that is in
mathematical statistics has yet to have
been derived as fully clean applied mathematics
from only something like Neveu. Still, basically
mathematical statistics is applied probability
which in principle can all be done back to
Neveu and only a little more, e.g., matrix theory, some
combinatorics, maybe some group theory for
bootstrap and resampling plans, etc.
Again, my post was on the foundations of
probability and showing that there we did not
need to mention either frequentism or Bayesianism.
That is, maybe my post would be helpful
for people struggling with frequentism or
Bayesianism -- I'm saying, since 1933,
for just the foundations, get to
f'get about both of them.
I studied probability from a star student of
Cinlar at Princeton and from books by
Neveu, Chung, Breiman, Loeve, and others.
I did a lot in applied probability and
applied statistics. I've published in
mathematical statistics.
My Ph.D. dissertation was in applied
probability. So? I have
background enough to post.
But here I only wanted to
explain the foundations of probability
and not compete with anyone or compare
my expertise with anyone. Instead, I'm
just reporting some news current as of
1933.
It does appear to me that now
for nearly all serious work in
probability and stochastic processes,
the Kolmogorov foundations are
nearly universally accepted; thus,
readers are safe in taking seriously
the news I reported.
Maybe someday I will return to
pure and applied probability, but for
now my interests are in my startup.
There, now, mostly the work is in
software and other parts of business.
At the core, my startup is some
work in applied probability, but I did
that months ago and long since
have had the corresponding computations
in solid software.
So, for now, more background in
probability is not on my TODO list.
Again, here I'm just giving some
HN readers the news that as of
1933 get to f'get about the
frequentism and Bayesianism foundations
of probability.
Leaving aside that Kolmogorov's measure-theoretic approach is not the only axiomatic definition of probability (Cox's axioms yield a quite similar foundation, though with finite additivity only), you won't find anyone here that says there is a problem with the mathematical construction.
The frequentist/Bayesian debate is related to the INTERPRETATION of probability. Kolmogorov won't help you to map the real world to the probability space.
Let's say we have a loaded coin, we want to estimate the probability of getting tails (assume this is a i.i.d. random variable).
Alice decides to keep throwing until she gets a tail: she gets the sequence HHT
Bob decides to throw the coin three times: he gets the sequence HHT
Alice takes her event A, her sigma-algebra, the whole shebang, and produces an interval estimate for p.
Bob takes his event A (which happened to be the same) and his probability space (which is different because the experimental design is different), and produces a different estimate for p.
If you think that getting different results from the same data makes sense, you might be a frequentist.
If you think that it doesn't, you might be a Bayesian.
If you think that the question is not relevant because you can't derive the answer from your axioms, you might at least understand what we're talking about.
This is all mixed up. Alice has real random
variables X_1 (borrowing TeX notation for a subscript),
X_2, X_3. We assume that {X_i|i = 1, 2, 3} is independent.
We assume that for some number p in [0,1], P(X_i = 0) = p --
Alice uses 0 for H and 1 for T.
Alice observes that X_1 = 0, X_2 = 0, and X_3 = 1.
Now, for the set of real numbers, R, Alice has a
Borel measurable
function f: R^3 --> [0,1] and lets f(X_1, X_2, X_3) be
her estimate of p.
Okay, if that is what you meant.
Bob does much the same with real random variables
Y_i, i = 1, 2, 3 and function g: R^3 --> [0,1]
and lets his estimate of p be g(Y_1, Y_2, Y_3).
And Bob observes Y_1 = 0, Y_2 = 0, and Y_3 = 1.
If functions f and g are the same, then, in this
case, that is, with the data Alice and Bob
observed, Alice and
Bob get the same estimate for p. Else
if f and g are different, then, even if
the data they observe is the same, they
might
get different estimates. Even if f = g,
since each of Alice and Bob is
flipping the coin for themselves,
they need not get the same estimate for p.
Note: Since f is Borel measurable, we can
set real random variable Z = f(X_1, X_2, X_3) and
ask for E[Z], etc. Sometimes this step is useful.
That is, our estimator of p is also a random variable.
We might like to have E[Z] = p; in this case, Z is
an 'unbiased' estimator of p.
Also we can ask for the variance of Z, say, Var(Z), and
maybe we want Var(Z) to be small. If we can show
that Var(Z) is the smallest among all
Borel measurable f, then the choice Alice made
for f is a 'minimum variance' estimator.
If Alice gets charged money for being wrong, then
we can try to minimize the expected value of
what Alice gets charged, and here we have
a case of 'statistical decision theory'.
Since order statistics are always sufficient,
with the assumptions we have, Alice need
only be told 2 zeros and 1 1 and can
f'get about the rest.
I see no surprises or difficulties here.
But the sample space Omega is the same for
both Alice and Bob. And, for both Alice
and Bob, there is only one
trial, that is, only one point little omega,
in Omega involved. That is, in more detail,
X_1 is a function, the function
X_1: Omega --> R, and in our case for
our trial little omega we have
X_1(little omega) = 0. That is, usually in
the notation we suppress little omega. Alice
and Bob are both using the same trial little omega
and the same sigma algebra script F on the same
sample space Omega.
We have no reason to believe that random variables
X_1: Omega --> R and Y_1: Omega --> R are equal.
Thus, given a Borel subset K of R, the events
X_1^{-1}(K) in the sigma algebra script F on Omega
and Y_1^{-1}(K) also in the sigma algebra script
F on Omega need not be the same.
That's a little of how applied probability
based on 'modern probability' works.
I see no problems and no need to consider
frequentism or Bayesianism.
"We might like to have E[Z]=p","maybe we want Var[Z] to be small"... How do you select f and g? That's the difficulty. You see no problems because you're happy playing with your Borel measurable thingies. Try to go further in the 'statistical decision' field and see how longer can you avoid frequentist considerations.
What do you think of the "likelihood principle"? Alice and Bob get the same data and the same likelihood function. Should they make the same inference?
Do you think unbiasedness is an important property for an estimator? What makes an estimator admissible? Do you see a problem with an estimator that provides negative values for non-negative variables?
By the way, I don't think you analysis is correct: Alice doesn't have random variables X_1, X_2, X_3. The possible events for her are T,HT,HHT,HHHT,HHHHT,... (she stops when she get tails, but not before).
EDIT: In case it's not yet clear: saying that "our estimator of p is also a random variable" and looking at its sampling distribution is frequentism.
> How do you select f and g? That's the difficulty.
You see no problems because you're happy playing
with your Borel measurable thingies.
Borel measurablility is important work; it sounds
like you have contempt for it. There is no need or
justification for contempt.
We should mention Borel measurability, as I did, if
we are carefully considering the Kolmogorov
foundations, but in practice Borel measurability
means essentially nothing since cooking up a
function, such as f or g in what I wrote, that is
not Borel measurable is so tricky that essentially
any function anyone would select for f or g will be
Borel measurable. The usual example of a function
not Borel or Lebesgue measurable uses the axiom of
choice -- we're talking tricky stuff won't see in
SPSS, SAS, R, Mathematica, Matlab, big data, machine
learning, etc.
For picking f or g, let's see, from our definition
of p above,
is an unbiased estimator of p. We expected something
else? This was difficult?
Here we specified a function g and showed that it
gives an unbiased estimator of p and did this
without mentioning the trial little omega, the
sample space big Omega, or the sigma algebra of
events script F. So, what we did is just an
elementary part of standard junior level
introductory mathematical statistics, and so would
be responses to your other questions.
And, again, we never mentioned frequentism or
Bayesianism. Yet again, when considering the
mathematical foundations of probability, I see no
need to consider either frequentism or Bayesianism.
Again, the Kolmogorov foundations work just fine for
long standard probability, applied probability,
mathematical statistics, and applied statistics. No
worries.
For Alice, a different derivation is required.
My point is that from Kolmogorov we have some rock
solid foundations for probability and, then, do not
have to consider either frequentism or Bayesianism.
So, students struggling over frequentism or
Bayesianism can relax and just f'get about these
two.
> So, for Bob's case, (....) so that our estimator Z = g(X_1, X_2, X_3) = 1 - (1/3) (X_1 + X_2 + X_3) is an unbiased estimator of p. We expected something else? This was difficult? (....) For Alice, a different derivation is required.
A different derivation that you tried (I saw your comment appear briefly). You proposed Z=1/N, which is (as before) the maximum-likelihood estimator. And the likelihood function is the same for Alice and for Bob, so it is not surprising that we get the same estimate p=1/3 in both cases. But after "proving" that "again our estimator Z is unbiased" I guess you noticed that this was not in fact correct. Did you expect something else? What was the difficulty?
> My point is that from Kolmogorov we have some rock solid foundations for probability and, then, do not have to consider either frequentism or Bayesianism.
Both are based on probability, how is probability going to replace them?
> So, students struggling over frequentism or Bayesianism can relax and just f'get about these two.
Sure, they can avoid thinking about the different approaches to inference... and just do it in the frequentist way. Assuming that the unknown parameter is fixed, that the estimator is a random variable, that the criteria to select an estimator are the unbiasedness, consistency or asymptotic distribution... all of these are frequentist considerations (even if you say that you "do not have to consider either frequentism or Bayesianism").
You never answered my questions about Alice and Bob getting different interval estimates (even though the outcome of their experiments is identical and their model for the loaded coin as well). We've seen that they agree on their (MLE) point estimate, but their confidence intervals will be different.
I imagine you accept the frequentist idea of confidence intervals, and agree that they will be different because the distribution of potential experiment outcomes is different.
Do you think that Alice and Bob can get different conclusions from the same model and the same data?
Carol performs another experiment. She starts by rolling a die to see if she'll do it like Alice (even) or like Bob (odd). So with 50% probability she will throw the loaded coin 3 times and with 50% probability she will do it until she gets a tail. She gets 5 on the die, she throws the loaded coin three times and gets 'HTH'. How will you select your estimator for p? Do you use the results you obtained for Bob? Do you repeat you analysis considering the mixture of both experiments?
That's a contribution to this
discussion about the foundations of
probability, pure math, applied math,
probability theory, statistics,
econometrics, or just an insult?
If your point is that all branches of statistics are built on the same axiomatic foundations of probability, of course I agree.
> To respond to your question about sample size, I'd have to look into some of the details of your question. And I can say now, that question, those details, and any answer all have essentially nothing to do with the axiomatic foundation of probability I gave; the answer is the same or essentially so independent of those foundations down in the deep sub basement of the subject.