Part of being a conscientious scholar is admitting one's earlier mistakes. While it can be tempting to toss earlier mistakes into a memory hole and hope nobody notices them, it's something I try to avoid. Another temptation can be retconning what one said so as to imply that one meant something that one can still maintain. I don't suffer from the latter temptation, but I do sometimes experience the former temptation. I try hard to resist both.
Let it be said here at the outset that the correction here does not mean that "undesigned coincidences are not a thing," "undesigned coincidences have been refuted," "professional publication on undesigned coincidences is fraudulent and needs to be retracted" or anything of the kind. Skeptics, I'm lookin' at you; if you try to use this post for that purpose, you're making a mistake. Also, even if you don't slog through this whole post, please read the last paragraph.
In 2020 an article (by me) on the probabilistic analysis of
undesigned coincidences was published in the high-level journal Erkenntnis. The accepted manuscript
version of that article can be found here, and it now has a dated cover sheet
with a short update note linking back, in turn to this blog post for further
details.
Recently I have been reviewing that article while preparing
for a talk on the mathematics of undesigned coincidences. I've realized that in
the appendix to that article, where I give some illustrative worked examples,
Example 3 has a couple of errors. Here are the errors and the corrected work:
The first (and greater) error is that, when discussing
dependence under H, I treat the fact that, when we are considering the
conditional probability given H, we are taking H to be true, as if it confers
additional reason to take the witness to be reliable. This is incorrect, since
this is a mere conditional probability and does not actually involve receiving
further independent evidence that H is in fact true.
The second error is a sheer matter of incorrect calculation
using the model. While I intended (and even emphasized) that the subhypothesis
of H that I dubbed H', and its negation, should screen off testimony 1 from
testimony 2, under the assumption of H, I did not in fact build that screening
off into the model correctly. This error was related to my not properly taking
account of the fact that, in this example, I am creating dependence under H by
stipulating that a portion of the probability space under H actually gives each
of the individual items of testimony higher probability than the rest of the
space under H. This differs from the nature of the dependence under ~H modeled
in the previous two examples; the relevant subhypotheses there directly
hypothesize collusion or copying and make no difference to the probability of
an individual item.
Here is a corrected worked example, modeling a dovetailing
undesigned coincidence:
H is the umbrella hypothesis that some event took place. (In
the text of the article the event is a picnic.)
Let H' be the subhypothesis that the event took place and
also that some further fact about the event is true (in the paper, this was
that a picnic took place and also that a Joyce scholar was present at the
picnic).
Let P(H'|H) = .05
Suppose that what each of the witnesses says confirms not
just H but also H'.
Contrary to what I did in the published paper, the
credibility of the witnesses 1 and 2 should be kept at a Bayes factor of 10/1
(as it already was) for their information that confirms H', even though this is
a detail of their testimony that H, since no independent information (aside
from their testimonies) confirming H is being modeled.
For the sake of not complicating the model yet further,
suppose that each of their testimonies confirms H' to the same extent.
Both testimonies are screened off from one another given H'
and given its negation, modulo H, but they are not screened off from one
another by H itself, because they both confirm H'. The probability of the
conjunction of the testimonies is higher
given H' than given H and than given (H & ~H'), because the probability of
the individual items is higher given H' than given (H & ~H'). This produces
the fruitful dependence that actually creates a kind of "bonus"
probabilistic boost for H from the conjunction of these testimonies.
Let P(T2|H & ~H') = .02
And the same for T1.
Let P(T2|H & H') = .2
And the same for T1.
Therefore, the overall probability of T1 individually and T2
individually given H is
(.95)(.02) + (.05)(.2) = .019 + .01 = .029
(It might be possible to set all the probabilities just right
so that this comes out as .03 as it previously did in example 3 in the published paper, but I did not keep trying all possibilities, and
.029 was the closest I was able to get in the numbers I attempted, while
keeping everything else as it should be—for example, the BF between H' and ~H'
as 10/1. Since this is similar to the other worked examples where the
individual probability of each item of testimony given H is .03, and since this
is merely an illustrative example, I decided to go with this modeling. Nothing
rides on having the total individual probability given H be .03 rather than .029,
as long as everything else is done accurately within the assumptions stated.)
Let the individual, independent Bayes factors between H and
~H be 10/1 as in the other appendix examples, so that P(T1|H)/P(T1|~H) =
.029/.0029 and the same for T2, so that treated as independent the joint BF
would be 100/1, as in the other examples.
Confirmation of H' within H by a 10/1 internal Bayes factor:
Since P(T1|H & H') = .2 and P(T1|H & ~H') = .02,
under H, T1 confirms H' by a factor of 10/1, and so does T2.
By the odds form
P(H'|H)/P(~H'|H) x P(T1|H' & H)/P(T1|~H' & H) =
P(H'|T1 & H)/P(~H'|T1 & H)
.05/.95 x 10/1 = .5/.95
To turn these odds (under H, giving us the probability of H'
given either of the testimonies and H) into a probability,
.5/(.95 + .5) ≈ .3448
So given H and either one of the testimonies (say, T1), the
probability of subhypothesis H' is approximately .3448 and the probability of
its negation is approximately .6552.
We have stipulated that, under H, both H' and ~H' screen off
T1 from T2 and that they each have the same probabilities (.2 and .02
respectively) given (H & H') and (H & ~H'). Given such screening, the conditional probabilities of .2 and .02 remain the same.
Therefore, the probability of T2 given (H & T1) =
(.3448)(.2) + (.6552)(.02) = .06896 + .013104 = .082064
To calculate the fruitful dependence factor between T1 and
T2 under H:
P(T1|H) x P(T2|H & T1)/P(T1|H) x P(T2|H) = (by cross
canceling P(T1|H) on the top and bottom)
P(T2|H & T1)/P(T2|H) =
.082064/.029 ≈ 2.82979/1
This, then, is the correction factor by which T1 confirms T2
given H—which will be helpful to the confirmation given by the conjunction (T1
& T2) to H.
This is almost exactly half of the size of the dependence
factor of 5.617 given in the published paper in the worked Example 3, modeling
dovetailing details.
In the text of the paper other than the appendix, the
closest that I come to stating the erroneous extra boost of the witness's
credibility is this sentence on p. 21:
Given that P (asserted by Source 1)
is true, there is some reason to
think that Source 1 is truthfully reporting about that day, including the extra
detail about the Joyce scholar's presence.
This should be reworded along these lines:
Since Source 1 has, as stipulated,
some credibility for what he attests, this credibility can be applied to the
extra detail about the Joyce scholar's presence.
Concerning the correction factor for dependence under ~H, I
argued in the appendix section on Example 3 that dependence on ~H (which
weakens the case for H) should not be any higher
than in example 2 (varied but not dovetailing details), where it is 2.33 to 1
against H, and plausibly should be lower.
I stand by that. In that calculation, the only dependence
arose from modeling two subhypotheses under ~H of PC (partial/possible copying)
and Cr (crafty copying) as having a relative advantage in predicting the
conjunction (T1 & T2) over complete independence modulo ~H. If that
relative advantage is kept proportionally
identical in Example 3 (the one I am recalculating), then of course that
dependence factor will be exactly the same.
If the dependence given ~H is kept the same as in example 2,
on the new calculation the approximately 2.82979 dependence given H more than
offsets the approximately 2.33 dependence given ~H, though obviously not as
much as the dependence factor of twice that much in the worked example in the
paper.
I then pointed out that arguably in the dovetailing case hyper-crafty collusion/copying would be
necessary so as to produce the dovetailing between the details with the
appearance of casualness/undesignedness. In other words, partial/possible
copying would do an even worse job at uniting the testimonies, as would crafty
copying. (Full copying of course is just out of the picture here, as it is in example
2.) I suggested that we should even reasonably consider eliminating a
dependence factor under ~H altogether in the case of dovetailing details, since
in order to have T1 and T2 as described here we have to have not just variation
but extra craftiness to think of varying in a way that is both unobvious and
dovetailing, which might simply be overlooked by an audience and which
colluders are unlikely to think of. And it's even less likely than
non-dovetailing variation to come about by partially copying the content while
also varying in some unthinking fashion.
As stated in the paper, if dependence given ~H is eliminated
due to these considerations, then the dependence given H is all gain for the
confirmation of H from the conjunction of the testimonies, over and above the
100/1 Bayes factor treating them as independent.
It is also worth emphasizing that these numbers are merely illustrative. If the two witnesses have even higher credibility in the first place, that will affect this type of example particularly strongly, since the witness's existing credibility should apply to the subhypothesis H'.