Saturday, August 15, 2026

Correction of worked example of dependence factor in undesigned coincidence article

Part of being a conscientious scholar is admitting one's earlier mistakes. While it can be tempting to toss earlier mistakes into a memory hole and hope nobody notices them, it's something I try to avoid. Another temptation can be retconning what one said so as to imply that one meant something that one can still maintain. I don't suffer from the latter temptation, but I do sometimes experience the former temptation. I try hard to resist both.

Let it be said here at the outset that the correction here does not mean that "undesigned coincidences are not a thing," "undesigned coincidences have been refuted," "professional publication on undesigned coincidences is fraudulent and needs to be retracted" or anything of the kind. Skeptics, I'm lookin' at you; if you try to use this post for that purpose, you're making a mistake. Also, even if you don't slog through this whole post, please read the last paragraph.

In 2020 an article (by me) on the probabilistic analysis of undesigned coincidences was published in the high-level journal Erkenntnis. The accepted manuscript version of that article can be found here, and it now has a dated cover sheet with a short update note linking back, in turn to this blog post for further details.

Recently I have been reviewing that article while preparing for a talk on the mathematics of undesigned coincidences. I've realized that in the appendix to that article, where I give some illustrative worked examples, Example 3 has a couple of errors. Here are the errors and the corrected work:

The first (and greater) error is that, when discussing dependence under H, I treat the fact that, when we are considering the conditional probability given H, we are taking H to be true, as if it confers additional reason to take the witness to be reliable. This is incorrect, since this is a mere conditional probability and does not actually involve receiving further independent evidence that H is in fact true.

The second error is a sheer matter of incorrect calculation using the model. While I intended (and even emphasized) that the subhypothesis of H that I dubbed H', and its negation, should screen off testimony 1 from testimony 2, under the assumption of H, I did not in fact build that screening off into the model correctly. This error was related to my not properly taking account of the fact that, in this example, I am creating dependence under H by stipulating that a portion of the probability space under H actually gives each of the individual items of testimony higher probability than the rest of the space under H. This differs from the nature of the dependence under ~H modeled in the previous two examples; the relevant subhypotheses there directly hypothesize collusion or copying and make no difference to the probability of an individual item.

Here is a corrected worked example, modeling a dovetailing undesigned coincidence:

H is the umbrella hypothesis that some event took place. (In the text of the article the event is a picnic.)

Let H' be the subhypothesis that the event took place and also that some further fact about the event is true (in the paper, this was that a picnic took place and also that a Joyce scholar was present at the picnic).

Let P(H'|H) = .05

Suppose that what each of the witnesses says confirms not just H but also H'.

Contrary to what I did in the published paper, the credibility of the witnesses 1 and 2 should be kept at a Bayes factor of 10/1 (as it already was) for their information that confirms H', even though this is a detail of their testimony that H, since no independent information (aside from their testimonies) confirming H is being modeled.

For the sake of not complicating the model yet further, suppose that each of their testimonies confirms H' to the same extent.

Both testimonies are screened off from one another given H' and given its negation, modulo H, but they are not screened off from one another by H itself, because they both confirm H'. The probability of the conjunction of the testimonies is higher given H' than given H and than given (H & ~H'), because the probability of the individual items is higher given H' than given (H & ~H'). This produces the fruitful dependence that actually creates a kind of "bonus" probabilistic boost for H from the conjunction of these testimonies.

Let P(T2|H & ~H') = .02

And the same for T1.

Let P(T2|H & H') = .2

And the same for T1.

Therefore, the overall probability of T1 individually and T2 individually given H is

(.95)(.02) + (.05)(.2) = .019 + .01 = .029

(It might be possible to set all the probabilities just right so that this comes out as .03 as it previously did in example 3 in the published paper, but I did not keep trying all possibilities, and .029 was the closest I was able to get in the numbers I attempted, while keeping everything else as it should be—for example, the BF between H' and ~H' as 10/1. Since this is similar to the other worked examples where the individual probability of each item of testimony given H is .03, and since this is merely an illustrative example, I decided to go with this modeling. Nothing rides on having the total individual probability given H be .03 rather than .029, as long as everything else is done accurately within the assumptions stated.)

Let the individual, independent Bayes factors between H and ~H be 10/1 as in the other appendix examples, so that P(T1|H)/P(T1|~H) = .029/.0029 and the same for T2, so that treated as independent the joint BF would be 100/1, as in the other examples.

Confirmation of H' within H by a 10/1 internal Bayes factor:

Since P(T1|H & H') = .2 and P(T1|H & ~H') = .02, under H, T1 confirms H' by a factor of 10/1, and so does T2.

By the odds form

P(H'|H)/P(~H'|H) x P(T1|H' & H)/P(T1|~H' & H) = P(H'|T1 & H)/P(~H'|T1 & H)

.05/.95 x 10/1 = .5/.95

To turn these odds (under H, giving us the probability of H' given either of the testimonies and H) into a probability,

.5/(.95 + .5) ≈ .3448

So given H and either one of the testimonies (say, T1), the probability of subhypothesis H' is approximately .3448 and the probability of its negation is approximately .6552.

We have stipulated that, under H, both H' and ~H' screen off T1 from T2 and that they each have the same probabilities (.2 and .02 respectively) given (H & H') and (H & ~H'). Given such screening, the conditional probabilities of .2 and .02 remain the same.

 Therefore, the probability of T2 given (H & T1) =

 (.3448)(.2) + (.6552)(.02) = .06896 + .013104 = .082064

 To calculate the fruitful dependence factor between T1 and T2 under H:

 P(T1|H) x P(T2|H & T1)/P(T1|H) x P(T2|H) = (by cross canceling P(T1|H) on the top and bottom)

 P(T2|H & T1)/P(T2|H) =

 .082064/.029 ≈ 2.82979/1

 This, then, is the correction factor by which T1 confirms T2 given H—which will be helpful to the confirmation given by the conjunction (T1 & T2) to H.

This is almost exactly half of the size of the dependence factor of 5.617 given in the published paper in the worked Example 3, modeling dovetailing details.

In the text of the paper other than the appendix, the closest that I come to stating the erroneous extra boost of the witness's credibility is this sentence on p. 21:

Given that P (asserted by Source 1) is true, there is some reason to think that Source 1 is truthfully reporting about that day, including the extra detail about the Joyce scholar's presence.

 This should be reworded along these lines:

 Since Source 1 has, as stipulated, some credibility for what he attests, this credibility can be applied to the extra detail about the Joyce scholar's presence.

Concerning the correction factor for dependence under ~H, I argued in the appendix section on Example 3 that dependence on ~H (which weakens the case for H) should not be any higher than in example 2 (varied but not dovetailing details), where it is 2.33 to 1 against H, and plausibly should be lower.

I stand by that. In that calculation, the only dependence arose from modeling two subhypotheses under ~H of PC (partial/possible copying) and Cr (crafty copying) as having a relative advantage in predicting the conjunction (T1 & T2) over complete independence modulo ~H. If that relative advantage is kept proportionally identical in Example 3 (the one I am recalculating), then of course that dependence factor will be exactly the same.

If the dependence given ~H is kept the same as in example 2, on the new calculation the approximately 2.82979 dependence given H more than offsets the approximately 2.33 dependence given ~H, though obviously not as much as the dependence factor of twice that much in the worked example in the paper.

I then pointed out that arguably in the dovetailing case hyper-crafty collusion/copying would be necessary so as to produce the dovetailing between the details with the appearance of casualness/undesignedness. In other words, partial/possible copying would do an even worse job at uniting the testimonies, as would crafty copying. (Full copying of course is just out of the picture here, as it is in example 2.) I suggested that we should even reasonably consider eliminating a dependence factor under ~H altogether in the case of dovetailing details, since in order to have T1 and T2 as described here we have to have not just variation but extra craftiness to think of varying in a way that is both unobvious and dovetailing, which might simply be overlooked by an audience and which colluders are unlikely to think of. And it's even less likely than non-dovetailing variation to come about by partially copying the content while also varying in some unthinking fashion.

As stated in the paper, if dependence given ~H is eliminated due to these considerations, then the dependence given H is all gain for the confirmation of H from the conjunction of the testimonies, over and above the 100/1 Bayes factor treating them as independent.

It is also worth emphasizing that these numbers are merely illustrative. If the two witnesses have even higher credibility in the first place, that will affect this type of example particularly strongly, since the witness's existing credibility should apply to the subhypothesis H'.

No comments: