Saturday, October 5, 2013

The Misleading Krauss? (p. 2)

Krauss and Craig
< Prev   |   1   |   2   |   3   | 

My interaction with Krauss



There’s more to the story too. Craig made a response at Reasonable Faith arguing that Krauss’s use of the edited email was misleading, and Krauss made a post on Facebook that was appears to be a response to Craig’s claim about the use of the edited email being misleading. To quote Krauss’s Facebook post entry:
From me and Alex Vilenkin--sigh--in muted response to some claims that have been posted by some whose buttons have probably been pushed by being wrong:

"In response to the noise regarding the use of an email communication between the two of us in a dialogue with William Lane Craig, there are two relevant points we have decided to make.

1. we both willingly agreed to the request from Dr. Craig to have the full email, which had been edited on the powerpoint slide simply to save time during a 15 minute presentation by Krauss, as there was nothing in the full correspondence that either of us were concerned about sharing.

2. we both agree that the edited version does not distort the content or ideas expressed in the original email at all. Those who are claiming otherwise, including apparently Dr. Craig, are mistaken.

Lawrence Krauss and Alex Vilenkin"
Last week I had a little electronic interaction with Krauss himself. Here was my Facebook post (with the text formatted a bit for neatness).
@Lawrence Krauss

I think you ought to realize why some people might find the edited version of the email potentially misleading; certainly I would have been given a false impression if I hadn’t known a bit of Vilenkin’s background. To see why, let’s consider a concrete example:
Any theorem is only as good as its assumptions. The BGV theorem says that if the universe is on average expanding along a given worldline, this worldline cannot be infinite to the past.

A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. . . . . .
Whether you realize it or not, this sort of thing might give one the impression that the BGV theorem doesn’t provide significant evidence for there being a beginning of the universe due to the loophole mentioned here. But now look at it in context:
Any theorem is only as good as its assumptions. The BGV theorem says that if the universe is on average expanding along a given worldline, this worldline cannot be infinite to the past.

A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. They had to assume though that the minimum of entropy was reached at the bounce and offered no mechanism to enforce this condition. It seems to me that it is essentially equivalent to a beginning.
Look at this way: a person reading the edited email might say to a theist, “See! There’s a loophole here to get around the supposed evidential force the BGV theorem has for a beginning of the universe.” But when one reads the quote in context, a somewhat different impression is given. By leaving the “essentially equivalent to a beginning” etc. part out, the edited email delivers a false impression about the nature of the BGV theorem and the beginning of the universe, even if that wasn’t your intention.

Now consider another example:
. . . Jaume Garriga and I are now exploring a picture of the multiverse where the BGV theorem may not apply. In bubbles of negative vacuum energy, expansion is followed by contraction. . . However, it is conceivable (and many people think likely) that singularities will be resolved in the theory of quantum gravity, so the internal collapse of the bubbles will be followed by an expansion. In this scenario, . . . it is not at all clear that the BGV assumption (expansion on average) will be satisfied.
Whether you realize this or not, this kind of gives the impression that Vilenkin might believe the multiverse avoids the evidential force of the BGV with respect to a beginning of the universe. Think of an atheist looking at this quote and using it against a theist saying, “See! Vilenkin doesn’t think the BGV theorem provides much evidence for a beginning of the universe.” But including the omitted parts would deliver a somewhat different impression:
On the other hand, Jaume Garriga and I are now exploring a picture of the multiverse where the BGV theorem may not apply. In bubbles of negative vacuum energy, expansion is followed by cocntraction, and it is usually assumed that this ends in a big crunch singularity. However, it is conceivable (and many people think likely) that singularities will be resolved in the theory of quantum gravity, so the internal collapse of the bubbles will be followed by an expansion. In this scenario, a typical worldline will go through a succession of expanding and contracting regions, and it is not at all clear that the BGV assumption (expansion on average) will be satisfied. I suspect that the theorem can be extended to this case, maybe with some additional assumptions.
So there are a couple key omissions here: the sort of model which avoids a big crunch singularity is usually assumed to be incorrect, and Vilenkin suspects the theorem can be extended to this case. Omitting all this has at least the potential to deliver a false impression with respect to what Vilenkin believes about the evidential force the BGV theorem has for the beginning of the universe.

It seems to me that William Lane Craig asked some valid questions (http://www.reasonablefaith.org/...):
Why didn’t Krauss read the sentence, “It seems to me that it is essentially equivalent to a beginning”? Because it was too technical? Is this the transparency, honesty, and forthrightness that Krauss extols? (By the way, Vilenkin’s criticism of these models is the same one that Vilenkin makes in his Cambridge paper: far from showing an eternal past, these models actually feature a universe with a common beginning point for two arrows of time.)

And why did Krauss delete Vilenkin’s caveat that the BGV theorem can, in his estimation, be extended to cover the case of an expanding and contracting model such as Garriga and Vilenkin are exploring? And why delete the remark that such a model is usually assumed to be incorrect?
Maybe there are valid answers to these questions, but the fact that these are pretty good questions to ask reveals that you might not realize how potentially misleading these omissions were when it came to the topic of the BGV theorem and the beginning of the universe. Again, maybe this wasn’t your intention, but omitting relevant details like this isn’t recommended if the goal is to prevent giving a false impression.
Krauss responded:
Wade thanks for your note.. however the points I made were, as I explained to Craig at the time: (1) it is possible that we live in a bouncing universe, even if I would suggest current evidence and theory make that less likely than the alternative. (2) BGV rely on extrapolating back to singularity.. quantum gravity could change everything and we don’t have a theory.. therefore at this point anything goes.. all points in Alex’s email as I showed it.. and beyond that, as I explained.. the universe having a beginning or not (especially poorly defined in the case that time arises after the big bang) is irrelevant anyway, and he should stop fixating on it.. Moreover, as you will see, it was one slide out of 15 or so in a 15 min presentation, and the issues regarding entropy are interesting, but not appropriate for a general, non-scientific audience discussion with craig.
Many of the points I brought up he did not address; indeed there was an almost complete lack of engagement with why the quote was potentially misleading (e.g. Krauss did not answer Craig’s question about why “It seems to me that it is essentially equivalent to a beginning” was omitted). His point about the universe having a beginning or not being irrelevant was puzzling as well as clearly false, since the beginning of the universe is a obviously crucial part of Craig’s argument. Krauss doesn’t address the “It seems to me that it is essentially equivalent to a beginning” sentence and instead continues to talk about entropy as if the “It seems to me…” sentence didn’t exist, just as he did with Craig earlier.

At that point I realized it was probably hopeless to convince Krauss that his edited email was potentially misleading, so I gave him some advice on how to better attack the argument. After all, if he agrees with Craig that the “universe begins to exist” premise is more plausible than its denial, it would be better to attack some other part of the argument. It was my modest hope that my advice would help lead to less misleading statements and more substantive objections.

It’s easy to see how some would suspect Krauss of intentionally misleading his audience with the edited email, e.g. leaving out the “essentially equivalent to a beginning” part, but I don’t think that’s what quite happened. But then how to explain what Krauss did?

< Prev   |   1   |   2   |   3   | 

The Misleading Lawrence Krauss?

Krauss and Craig
< Prev   |   1   |   |   3   | 

There’s been a bit of hullabaloo on the internet about whether noted atheist Lawrence Krauss took Alexander Vilenkin out of context in a public discussion with Christian apologist William Lane Craig. I electronically interacted with Krauss about this matter a little. But for those who aren’t aware of the matter, first a bit of background introducing how Krauss (arguably?) mislead people in a big event in Australia.

Brief Background



Something called the kalam cosmological argument (KCA) goes something like this:
  1. Anything that begins to exist has a cause.
  2. The universe begins to exist.
  3. Therefore, it has cause.
Further arguments are given to show that the cause of the universe is (among other things) a transcendent personal cause. If we have adequate grounds for thinking the universe has a transcendent personal cause, this gives at least some evidence for the truth of theism. You can see justification for premise 1 at my Anything that begins to exist has a cause article.

What about the second premise? Craig argues that the scientific evidence currently makes the universe having a beginning more probable than not, and that the Borde-Guth-Vilenkin (BGV) theorem plays a significant role in the scientific case for a cosmic beginning. Why? The BGV theorem says that any universe that has on average been expanding throughout its history cannot have an infinite past but must have a beginning. This would close off a lot of avenues for those who want to avoid a beginning of the universe, for the theorem holds even if our universe is part of a multiverse, the multiverse would also require a beginning if the “has on average been expanding throughout its history” part holds true.

In a nutshell, here’s what happened: Krauss quoted an edited email of Vilenkin that gave some caveats with respect to the efficacy of the Borde-Guth-Vilenkin (BGV) theorem. Theists like William Lane Craig have often used the BGV theorem in arguing for a beginning of the universe, whereas Krauss denoted some of the limits of the theorem in their discussion in Sydney Australia.

Misleading Editing?



You can see and hear Krauss mention the edited email at around 41:23 of the debate:



In the clip above, Krauss says that the theorem doesn’t hold true in the multiverse, but this misleading. Whether it holds true in the multiverse depends on whether the “has on average been expanding throughout its history” condition is met. If the condition is met, then the theorem does hold for the multiverse, and Krauss doesn’t mention this. Krauss also says that the theorem doesn’t work when there’s quantum gravity, which also isn’t necessarily true (more on that later). For those who want to read the edited email:
Hi Lawrence,

Any theorem is only as good as its assumptions. The BGV theorem says that if the universe is on average expanding along a given worldline, this worldline cannot be infinite to the past.

A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. ……

. . . Jaume Garriga and I are now exploring a picture of the multiverse where the BGV theorem may not apply. In bubbles of negative vacuum energy, expansion is followed by contraction... However, it is conceivable (and many people think likely) that singularities will be resolved in the theory of quantum gravity, so the internal collapse of the bubbles will be followed by an expansion. In this scenario, ... it is not at all clear that the BGV assumption (expansion on average) will be satisfied.

...of course there is no such thing as absolute certainty in science, especially in matters like the creation of the universe. Note for example that the BGV theorem uses a classical picture of spacetime. In the regime where gravity becomes essentially quantum, we may not even know the right questions to ask.
Craig wondered what came after the ellipsis of this claim:
A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. ……
At around 50:18 Krauss insinuates that Craig quotes Vilenkin incorrectly. With that in mind, notice the interaction between Krauss and Craig here:



For those who prefer reading a transcript, here’s a key part of what transpired in the segment above (I apologize in advance for not getting the transcript perfect below; I was hampered by the fact that Krauss repeatedly interrupted Craig and the moderator had only limited success in toning this down):
Craig: [Reading the edited email] The BGV theorem says that if the universe is on average expanding along a given worldline, this worldine cannot be extended [apparently realizing he misread it slightly] uh or cannot be infinite to the past.

A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. [Done reading email] Now the thing is Lawrence that in the very paper that I quoted from Alex Vilenkin last April, he shows specifically, by name, that the Aguirre & Gratton model, and the Carroll & Chen model, don’t work. That there—

Krauss: No no, he says that you have to make an assumption about entropy.

Craig: Yes, he—

Krauss: You have to make an assumption about the evolution of entropy at the point of minimum size.

Craig: Those—

Krauss: So you—you have to make an assumption which he would argue that they don’t have any rationale for. That’s not the same as saying that they’re wrong.

Craig: Well, yes it is. I—he—he argues that the—all of the evidence shows that the universe had a beginning and that in this model—

Krauss: I would agree that all the evidence shows—let’s look—let’s accept that fact.

Craig: All rught.

Krauss: All the evidence suggests our universe had a beginning.

Craig: Oh, OK—

Krauss: But we don’t KNOW! That’s what I keep telling you! Knowing and suspecting are two VASTLY different things!

Craig: I—I’m never saying that this is known with certainty. This is a—a mischaracterization on your part. What I argue is that the premise [the universe begins to exist] is more probably true than false. And actually you agree with me on that—

Krauss: Yeah but—

Craig: that the universe began to exist, but I want to focus—

Krauss: Well I say it’s likely that the universe began to exist—

Craig: —on this claim that I’ve somehow misrepresented Vilenkin—

Krauss: Well but the—but the key line is the last one. That the theorem breaks down—

Craig: Wait a minute.

Moderator: Hang on.

Krauss: — at the point it really matters.

Moderator: One at a time.

Craig: Now there’s more here. Because I noticed that at the end of this paragraph where the Carroll-Chen and Aguirre-Gratton models are mentioned that he specifically shows—

Krauss: [my guess is that Krauss might have thought that Craig was about to ask why Krauss omitted what came after the ellipsis] Because it was technical! He said—he talked about the fact that—

Moderator: Lawrence, hang on. Just let, just let him finish.

Craig: That he specifically shows that these models cannot be past eternal, uh and that they involve therefore a beginning of the universe, just like the others. I—

Krauss: You can do the math if you want.

Craig: Now wait—

Krauss: I’ll let you do it.

Craig: Now wait. I couldn’t help notice although it’s [the slide showing the edited email] is down from the screen now that there was a series of ellipses points following the paragraph—

Krauss: Yeah because it was technical! And I thought it was, you know—

Craig: Well I wonder what you deleted from the original letter. Could it—

Krauss: I—I just

Craig: Now wait—

Krauss: I just TOLD you!

Craig: Now wait. Yeah, but you didn’t—

Krauss: I just told you. He says that you assume of entropy at the lowest point for which there is no rationalization in the papers within the context of the model given, and therefore he finds them unpleasant.
A few specific points to notice: Krauss says the entropy thing is something Vilenkin finds “unpleasant.” Also, Krauss actually concedes that the universe likely had a beginning! This is notable (and surprising) since as Craig said the claim is that the “universe begins to exist” premise is more probably true than false.

To sum up a few key points of the above transcript though, Craig wonders what was after the ellipses points here:
Hi Lawrence,

Any theorem is only as good as its assumptions. The BGV theorem says that if the universe is on average expanding along a given worldline, this worldline cannot be infinite to the past.

A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. ……
Craig was also evidently trying to say that Vilenkin says that the models he mentioned don’t work with respect to extending them to an infinite past. Krauss assures Craig that the stuff after ellipses is technical. To be sure, there was some technical stuff after the ellipses, but there was also something else. Here’s a fuller part of the quote, with a key part emphasized by yours truly:
Any theorem is only as good as its assumptions. The BGV theorem says that if the universe is on average expanding along a given worldline, this worldline cannot be infinite to the past. A possible loophole is that there might be an epoch of contraction prior to the expansion. Models of this sort have been discussed by Aguirre & Gratton and by Carroll & Chen. They had to assume though that the minimum of entropy was reached at the bounce and offered no mechanism to enforce this condition. It seems to me that it is essentially equivalent to a beginning.
Perhaps Vilenkin does find the entropy assumption “unpleasant,” but he also says the last emphasized sentence in the quote above that Krauss never mentions. So why did Krauss omit the part about it being essentially equivalent to a beginning? Because it was too technical?

< Prev   |   1   |   |   3   | 

Friday, August 16, 2013

“If A, then probably C” entails “Probably, if A then C”

Application



The fact that “If A, then probably C” entails “Probably, if A then C” is useful in a lot of philosophical discussions. Take for example this argument:
  1. If atheism is true, then objective morality does not exist.
  2. Objective morality does exist.
  3. Therefore, atheism is false.
In Does Objective Morality Exist If God Does Not Exist? I basically argued that if atheism is in fact true, objective morality probably doesn’t exist, in which case premise (1) is probably true. I thus used an “If A, then probably C” claim to show that an “Probably, if A then C” claim is true.

I would’ve thought “If A, then probably C” entailing “Probably, if A then C” would be uncontroversial even among internet atheists, but in discussing an argument against the possibility of an infinite past, someone I dialogued with on Facebook claimed the following:
If A, then probably C

does not entail

Probably, If A then C.
If you encounter an internet atheist (or anyone else) who disputes that “If A, then probably C” entails “Probably, if A then C” you can point them to this article which features a mathematical proof demonstrating that “If A, then probably C” does indeed entail “Probably, if A then C.” Fortunately the proof requires nothing more difficult than high school (or middle school) mathematics. Don’t worry if your math is a bit rusty; I’ll give a crash course in some basic probability and set theory.


The General Idea



First, an explanation of what “If A, then C” means exactly. The “If A, then C” material conditional says it is not the case that A is true and C is false (this is often good enough for philosophical arguments, since in a true material conditional, when A is true, C is true as well—because a true material conditional prohibits C from being false when A is true).

The general idea is that “Given A, C is probably true” means it is probably not the case that A is true and C is false. While I think the general idea is somewhat intuitively obvious, this article will mathematically prove the general idea to be true.


Mathematical Background



If you’re already savvy in math (particularly with some basic set algebra and probability theory) feel free to skip this section and go straight to the proof. Otherwise I’ll introduce some basic math stuff so that folks who aren’t quite so math savvy can follow along.

Set Operations



To illustrate some set operations, suppose our “universe” consists entirely of natural numbers 1 through 9. Now let A and B be the following:
A = {1, 5, 9}
B = {1, 5, 7, 8}
C = {2, 3}
SymbolExampleExplanation
∈
(element of)
1 ∈ AFor any set S, x ∈ S means that x is an element of S.
∉
(not an element of)
1 ∉ CFor any set S, x ∉ S means that x is not an element of S.
∩
(intersection)
A ∩ B = {1, 5}Given sets S and T, S ∩ T contains all the elements x such that x ∈ S and x ∈ T.
∪
(union)
A ∪ B = {1, 5, 7, 8, 9}Given sets S and T, S ∪ T contains all the elements x such that x ∈ S or x ∈ T.
∅
(empty set)
A ∩ C = ∅The empty set is a set that doesn’t contain any members.
ξ
(universal set)
A ∪ A’ = ξ
B ∩ ξ = B
ξ is basically “everything” in whatever universe the sets are “talking about,” e.g. if we’re dealing with sets of lowercase alphabets, like {a, e, i, o, u}, the universal set would be the entire lowercase alphabet. Sometimes the universal set is depicted as U or S.
S’
(complement of S)
B’ = {2, 3, 4, 6, 9}The complement of set S, denoted as S or SC or S’ or −S (among other variants), are all the elements x such that x ∉ S and x ∈ ξ.


It should be remembered that the union (∪) is using the “inclusive-or,” and so A ∪ B would include all elements that are in both A and B.

One notable thing is how similar some set operations are to propositional logic:

Set TheoryRough Equivalent in Logic
A ∪ BA ∨ B (“A or B”)
A ∩ BA ∧ B (“A and B”)
A’, −A¬A (“not-A”), alternatively, ~A and −A


So for example, x ∈ (A ∪ B) means that x ∈ A or x ∈ B.

Set Algebra



There are certain equality rules with sets involving stuff like unions and complements. Here’s a sample of some algebraic set rules:

Commutative laws: A ∪ B = B ∪ A  |  A ∩ B = B ∩ A
Identity laws: A ∩ ξ = A  |  A ∪ ∅ = A  |  A ∩ ∅ = ∅
Complement laws: A ∪ A’ = ξ  |  A ∩ A’ = ∅  |  (A’)’ = A
Distributive laws: A ∪ (B ∩ C) = (A ∪ B) ∩ (A ∪ C)
A ∩ (B ∪ C) = (A ∩ B) ∪ (A ∩ C)


Probability Symbolism



Probability often uses the language of set theory to symbolize the probabilities of certain events happening. Here a set denotes an event, like getting three or higher when rolling a die, where an event is a set of one or more outcomes. So for example if we let F represent the event of “rolling a four or higher” for a six-sided die, the set of outcomes would look like this:
F = {4, 5, 6}
If we let T be the event of “getting a 3,” T would look like this:
T = {3}
F ∪ T symbolizes all the outcomes that are in F or T, which in this case is rolling a 3 or higher. Pr(F ∪ T) denotes the probability that the outcome will be a member of set F or T. Some basic probability symbolism:

Pr(A) = The probability of A being true; e.g. Pr(A) = 0.5 means “The probability of A being true is 50%.”
Pr(A|B) = The probability of A being true given that B is true. For example:
Pr(I am wet|It is raining) = 0.8
This means “The probability that I am wet given that it is raining is 80%.”
Pr(¬A) = The probability of A being being false (¬A is read as “not-A”); e.g. Pr(¬A) = 0.5 means “The probability of A being false is 50%.”
Pr(B ∪ C) = The probability that B or C (or both) are true.
Pr(B ∩ C) = The probability that B and C are both true.
Pr(A|B ∩ C) = The probability of A given that both B and C are true.


Some alternate forms:

One VersionAlternate Forms
Pr(A) P(A)
Pr(¬A)  Pr(~A), Pr(−A), Pr(AC)
Pr(B ∪ C) Pr(A ∨ B)
Pr(B ∩ C) Pr(B ∧ C), Pr(B&C)
Pr(A|B)Pr(A/B)


The alternate forms can be combined, e.g. an alternate form of Pr(H|E) is P(H/E).

Probability Rules



In addition to the mathematical symbolism, there are also a number of mathematical rules regarding probability. When events A and B have no outcomes in common, i.e. when A ∩ B =∅, events A and B are set to be mutually exclusive or disjoint. For example, “rolling a two or lower” and “rolling a five or higher” are mutually exclusive events for rolling a six-sided die. Two events are said to be independent of each other if the outcome of one does not affect the outcome of the other, e.g. rolling a 6 the first time and rolling a 5 the second time for a six-sided die. Because I think it makes things clearer in what I’ll do later in this article, I’ll use the symbolism ¬A to denote “not-A” rather than A’. With that in mind:

Rule NameRule
Addition rule:Pr(A ∪ B) = Pr(A) + Pr(B) when A ∩ B = ∅
General addition rule:Pr(A ∪ B) = Pr(A) + Pr(B) − Pr(A ∩ B), regardless of whether A ∩ B = ∅
(note that when A ∩ B = ∅, Pr(A ∩ B) = 0)
Complement rule:Pr(¬A) = 1 − Pr(A)
Multiplication rule:Pr(A ∩ B) = Pr(A) × Pr(B) when A and B are independent
General multiplication rule:Pr(A ∩ B) = Pr(A) × Pr(B|A), regardless of whether A and B are independent
(Pr(B|A) = Pr(B) when A and B are independent)


Notice that because of the general multiplication rule (and a bit of simple algebra), this is also true for any events A and B:
Pr(B|A) = 
Pr(A ∩ B)
Pr(A)
And that’s pretty much all the math background you’ll need to follow along.


The Proof



For this to work I’ll break the proof in separate steps. In math and logic, a lemma is a claim that is proved to demonstrate something else later in a proof. For this proof I’ll be using several lemmas.

Recall that the “If A, then C” material conditional means it is not the case that A is true and C is false. Thus the probability that the material conditional is true can be mathematically depicted as this:
Pr(¬(A ∩ ¬C)) = 1 − Pr(A ∩ ¬C)
Which means “The probability of it not being the case that A and ¬C are both true.”

By “If A, then probably C” I mean “Given A, C is probably true,” which in turn means that Pr(C|A) is high. So to show that high Pr(C|A) entails a high Pr(¬(A ∩ ¬C)), I want to prove the following:
Pr(C|A) ≤ 1 − Pr(A ∩ ¬C)
Or equivalently:
Pr(¬(A ∩ ¬C)) ≥ Pr(C|A)
Lemma (1): (A ∩ C) and (A ∩ ¬C) are disjoint (mutually exclusive). We can show that no element in the universe can be a member of both (A ∩ C) and (A ∩ ¬C). Let x be an arbitrary element and let’s suppose x is a member of both (A ∩ C) and (A ∩ ¬C). With a bit of math logic, we show that there can’t be any x such that x ∈ (A ∩ C) and x ∈ (A ∩ ¬C) by assuming there is such an x and deriving an impossibility, like so:
  1. x ∈ (A ∩ C) and x ∈ (A ∩ ¬C)
  2. (x ∈ A and x ∈ C) and (x ∈ A and x ∈ ¬C), from (1) and definition of ∩
  3. x ∈ A and x ∈ C and x ∈ A and x ∈ ¬C, from (2)
  4. x ∈ C and x ∈ ¬C, from (3)
Of course, it’s impossible for there to be an element that is a member of a set and its complement, since (C ∩ ¬C) = ∅. Thus (A ∩ C) and (A ∩ ¬C) are disjoint, i.e. (A ∩ C) ∩ (A ∩ ¬C) = ∅.

With this in mind, let ξ be the universal set.
A ∩ ξ = A
⇔ A ∩ (C ∪ ¬C) = A
⇔ (A ∩ C) ∪ (A ∩ ¬C) = A
Lemma (2): Since (A ∩ C) and (A ∩ ¬C) and are mutually exclusive, by the rules of probability:
Pr(A ∩ C) + Pr(A ∩ ¬C) = Pr(A)
With those two lemmas in mind, consider this statement:
Pr(C|A) = 
Pr(C ∩ A)
Pr(A)
Now we swap Pr(A) for Pr(A ∩ C) + Pr(A ∩ ¬C), and this is a legitimate move thanks to the equality proved in lemma (2):
Pr(C|A) = 
Pr(C ∩ A)
Pr(A ∩ C) + Pr(A ∩ ¬C)

 
⇔  Pr(C|A) = 
Pr(A ∩ C)
Pr(A ∩ C) + Pr(A ∩ ¬C)
Just to make this easier to read, let’s have x represent Pr(A ∩ C) like so:
Pr(C|A) = 
x
x + Pr(A ∩ ¬C)
Given some value for Pr(A ∩ ¬C), what is the highest Pr(C|A) possible? One hint is this: given some Pr(A ∩ ¬C), when x goes to zero, so does Pr(C|A); a smaller x means a smaller Pr(C|A).[1] So to get the highest Pr(C|A) value given some Pr(A ∩ ¬C), we want x to be as big as possible. Now since this is true:
Pr(A ∩ C) + Pr(A ∩ ¬C) = Pr(A) ≤ 1

    ⇔ Pr(A ∩ C) + Pr(A ∩ ¬C) ≤ 1

    ⇔ Pr(A ∩ C) ≤ 1 − Pr(A ∩ ¬C)
The highest Pr(A ∩ C) (and thus x) can be is 1 − Pr(A ∩ ¬C). So substituting the maximum value for x to obtain an upper limit for Pr(C|A) gives us this:
Pr(C|A) ≤ 
1 − Pr(A ∩ ¬C)
1 − Pr(A ∩ ¬C) + Pr(A ∩ ¬C)

 
⇔  Pr(C|A) ≤ 
1 − Pr(A ∩ ¬C)
1 + [−Pr(A ∩ ¬C)] + Pr(A ∩ ¬C)

 
⇔  Pr(C|A) ≤ 
1 − Pr(A ∩ ¬C)
1 + 0

 
⇔  Pr(C|A) ≤ 
1 − Pr(A ∩ ¬C)
1

 
⇔  Pr(C|A) ≤  1 − Pr(A ∩ ¬C)

 
⇔  Pr(C|A) ≤  Pr(¬(A ∩ ¬C))

 
⇔  Pr(¬(A ∩ ¬C)) ≥  Pr(C|A)
This means that Pr(¬(A ∩ ¬C)) must be at least as great as Pr(C|A), which means Pr(C|A) being high entails Pr(¬(A ∩ ¬C)) being high. This in turn means “If A, then probably C” entails “Probably, if A then C.”

In response, one could attack the relationship between “If A, then probably C” and Pr(C|A). But of course, Pr(C|A) is the probability of C given A. So Pr(C|A) being high means that given A, C is probably true. “If A, then probably C” is saying that given A, C is probably true. Hence, “If A, then probably C” entails a high Pr(C|A), which entails “Probably, if A then C.”





[1] We can prove this more rigorously with some simple calculus. Since Pr(A ∩ ¬C) is constant, we can replace Pr(A ∩ ¬C) in the equation below with k (to symbolize a constant) and take the derivative:
x
x + Pr(A ∩ ¬C)
The calculus would thus be this using the quotient rule:
d
dx
 
x
x + k
   = 
(x + k)(1) − (x)(1)
(x + k)²
 = 
x + k − x
x² + 2k + k²
 = 
k
x² + 2k + k²
Since both x and k symbolize possible probability values, they must each be in the interval [0, 1]. One can see that for all k > 0, all values of x ∈ [0,1] produce a positive value in the derivative above, which means the rate of change is always positive throughout the x ∈ [0,1] interval, which means the largest value in the [0,1] interval will be when x = 1. For when k = 0, the slope will be constant (not changing) for all x > 0. What about when x goes to 0? Obviously we can’t just plug in 0 for x when k = 0, but we can take the limit:
lim
x→0
 
0
x² + 2(0)² + 0²
 


 ⇔  lim
x→0
 
0
x²
 


We can then use L’Hôpital’s rule a couple times:

lim
x→0
 
0
x²
 


 ⇒  lim
x→0
 
0
2x
 


 ⇒  lim
x→0
 
0
2
   = 0
The slope is thus still constant (not changing) for all x ∈ [0,1] when k = 0, and the maximum value will still be found at x = 1.