Printed odds, where they exist
Some publishers state approximate odds on packaging or publish product breakdowns, like a stated ratio for a special rarity or a guaranteed slot structure per box. Where these exist they are the closest thing to ground truth, though even printed odds are typically approximate and can vary between print runs and regions. Reading the fine print on the product itself is genuinely one of the most underrated research techniques in the hobby.
Other publishers, including some of the largest, publish no pull rates at all. For their sets, every rate you have ever read was reverse engineered by the community, which is both impressive and a reason for humility about the numbers.
Community tallies, the citizen science of cardboard
Community estimates come from aggregating recorded openings, including creator box breaks, forum and spreadsheet projects where openers log results, and sites that collect structured tallies across many contributors. Tally enough packs and the observed frequency of each hit converges toward its true rate. This is the law of large numbers doing honest work on an extremely silly dataset.
The method has known biases worth respecting. People share exciting boxes more than boring ones, which can inflate apparent rates, a selection effect. Openings cluster around a set's launch window and may not represent later print runs. Definitions matter too, since one tracker's hit category may not match another's. Good projects document their methods, and great ones publish raw counts.
Sample size and confidence, or how sure is sure
Rare events need big samples. To estimate a roughly 1 in 100 rate with reasonable tightness, hundreds of packs is a sketch and thousands is a portrait. A tally of 200 packs that observed two hits is consistent with true rates from far below to far above 1 in 100, which is why a rate quoted from a small sample deserves wide mental error bars regardless of how many decimal places it flaunts.
This is why TCG Prayer attaches a sample size and a confidence label to rate data, and marks placeholder numbers as demo data. High confidence means a large, well documented sample. Low confidence means treat this as a rough sketch. Unknown means exactly that, and pretending otherwise would be the one kind of false prophecy we refuse to sell.
Why estimates differ between sources
Two honest sites can publish different rates for the same set because they tallied different samples, at different times, from different print waves and regions, using different category definitions. Neither is necessarily wrong. Rates can genuinely differ between production batches, and sampling noise does the rest. When sources disagree, prefer the bigger sample, the clearer methodology, and the more recent data, in roughly that order.
And whatever the number, remember what it is. An estimate of an average across enormous numbers of packs, not a schedule for your box, not a promise, and not something any ritual on this site can move. The data tells you the shape of the ocean. Your box is still one bucket of it.