A few people asked me about Nod's preprint about 2 spillovers. He...

He makes two arguments criticizing Pekar et al 2022.
A proper analysis of his 1st argument actually points in the opposite direction and strengthens Pekar's conclusions.
His 2nd argument is not well defined.
🧵
The simplest calculation would be you just simulate 2 epidemics, and count how often both form basal polytomies
That gets you to a bayes factor of somewhere around 4X, in favor of 2 spillovers, not 1:
x.com/tgof137/status…
Pekar required that the 2 polytomies both need to be equal and balanced, with each making 30-70% of the total genomes.
Pekar required this for the 1 introduction case but did not require it for the 2 introduction case.
If you randomly pick 2 points from that curve, the two aren't always going to be close to each other.
x.com/tgof137/status…
We're saying that it is possible for an epidemic that starts some random place in Wuhan to grow very slowly before it takes off.
That might represent something like, "one infected person from Yunnan visits Wuhan and starts the pandemic":
x.com/tgof137/status…
Some lab leakers think the market is the perfect place for Covid to spread, better than any other.
But we all agree it's a crowded building, so an introduction there quickly starts growing.
Suppose those are completely independent of each other, i.e. they happen in different locations.
I found it's about 9% odds that the two will become balanced at 30/70 or closer.
If I just divide 9% by 3%, I get bayes factor 3.
So, by that logic, Nod did find something, he reduced ~4 down to ~3.
How should we quantify that?
36/3 = 12, so the bayes factor is 12X in favor of 2 spillovers instead of 1.
But I'm glad he started this discussion and helped me understand that Pekar actually underestimated how likely 2 spillovers is!
Most of the people supporting the lab leak theory want to say that the market outbreak is meaningless, because markets are just really good at spreading Covid.
While simple math says the lab leak showing up at the raccoon dog shop is a 1 in 10,000 coincidence, they gave it 100% odds of happening.
You can make an argument that the market outbreak happening after a lab leak is higher, but if you do that, you can't also make Nod's argument.
There's no one cohesive lab leak theory that makes sense.
For the one introduction case, the virus needs to have 2 early mutations, and then both clades grow quickly to become balanced.
That's rare, only 3% likely.
But... should you actually model that diversity, and take 2 random samples from it?
In that case, Nod would say that Pekar should reduce the odds, because 2 mutations is possible, but it might be too many.
It's hard to take this seriously when Nod says that 2 mutations is too many and the other lab leakers say that 2 mutations is not enough.
I think that's what Nod is trying to do.
His writing is not particularly clear and I haven't read his code.
I think maybe what he's trying to do here is:
2. "simulate the growth and diversity in the number of animal cases as if that was a human epidemic"
3. "pick 2 animals at random for spillover"
I don't know. Did the market outbreak start with a single infected animal or many?
Do animal outbreaks grow via the same model as human outbreaks?
Pekar is simulating along a network of human social connections. Do raccoon dogs have the same network?
In that case, there were 2 spillovers from 2 Covid infected hamsters to people. The 2 spillovers had 5 mutations apart.
On average, it takes about 60 days to get 5 mutations.
We can thus conclude that there were 60,000 infected hamsters in that shop.
They gained viral diversity in prior transmission in warehouses, etc. They were mixed up, moved around, and eventually put in one shop. There were 2 shipments from the Netherlands to Hong Kong.
I expect that Nod will come back with more explanations for why these are actually genius ideas. And maybe I'm missing something here. But so far I'm not seeing much.
But it's mostly just increasingly obscure arguments to try to explain away the obvious bullseye on Huanan market.
x.com/tgof137/status…
"when did it start?"
"where did it start?"
"which MRCA?"
"how much time/how many cases before A/B split?"
And then compare that theory to real Wuhan data.
But instead he's just arguing that the 2 mutations happened immediately after spillover, by some unlikely random chance.
(Those numbers are not perfect, because I only had a few of these reversed clock epidemics to compute from -- these are extremely rare)
x.com/tgof137/status…
Is it unlikely that B would go to Huanan market, instead of any other place in Wuhan without suspicious animals?
(~1,000 vendors at Huanan out of 10 million people in Wuhan)
(You can bump that down, but then Nod is wrong)
x.com/tgof137/status…
Koopmans' map shows that a vendor at that shop had Covid on December 15th.
x.com/tgof137/status…
But maybe you can still keep other theories, by saying that market sample was deposited later.
And somewhere along the way, there were intermediates (even though none of the simulations had those)
It's usually the first case, not the latter.
If it does exist, it either dies out or grows to become very large:
x.com/tgof137/status…
So now the lab leak scenario is like:
3% for 2 polytomies *
< 10% for one clock reversal *
< 10% for second clock reversal *
< 1% for T/T intermediates that exist but don't grow
And he will be correct, I need to model both and calculate the ratios for all these things.
But I'm pretty confident that any reasonable model will still put the 2 spillovers hypothesis way ahead, and the A0 -> A -> B0 -> B scenario described in Lv et al 2024 will come out behind.









