Imagine two computers holding almost identical lists. Sending both lists in full would spend bandwidth on entries they already share. A compact summary of each set can make it possible to recover just the differences.
An invertible Bloom lookup table, or IBLT, is one data structure for doing this. It distributes information about each entry across several cells in a summary table. Comparing the two summaries cancels shared contributions, leaving information from which the differing entries can be recovered.
The table needs enough space for that recovery. The Self-sizing work estimates how many entries differ so that a later comparison can allocate an appropriate amount of space. Its estimator depends on the placement rule used to build the summary.
In a small reproduction I prepared for Self-sizing IBLT, two sets differed by exactly one entry. The code successfully recovered that entry. Its estimate of how many entries differed was about 0.678.
The interesting part was the rule behind that number. The estimator assumed that each entry contributed to three distinct cells. The mapper drew three positions and removed duplicates. For this input, its positions were [0, 16, 16], which became [0, 16]. One repeated draw had left the estimator and the mapper working with different assumptions.
I reported the mismatch with reproduction evidence. The original project subsequently corrected the mapper and credited my contribution in its public repository.
In the version I examined, that rule was supposed to give every entry three distinct cells. Three draws do not guarantee three distinct results.
A small example with the arithmetic exposed
The preserved example uses 64 cells and seed 0, which makes the position calculation reproducible. The input fingerprint is 0xa4712c74562914; a fingerprint is the numeric representation used here to place the entry.
After deduplication, only cells 0 and 16 receive it. Each has a count of one. The estimator forms a numerator from the squared cell counts, with a correction for their total:
2 − 4/64 = 1.9375.
Its denominator still assumes three distinct cells:
3 × (1 − 3/64) = 2.859375.
Dividing those quantities gives approximately 0.677596. With three distinct occupied cells, the numerator for this one-entry example would instead equal the denominator, giving an estimate of one.
This example isolates the mismatch without requiring a large benchmark. Recovery succeeded, so it would be misleading to describe it as lost data. It also does not refute a theorem whose three-cell assumption the implementation failed to meet. The frozen mapper code preserves the version behind the example.
Why the passing tests did not settle it
The existing tests and checks of the paper’s table values had passed in the preserved run records. But the calibration simulation retried a position when it encountered a duplicate, continuing until it had three distinct cells. The mapper in the reconciliation path stopped after removing duplicates.
The simulation followed the estimator’s assumption. A different execution path did not. Comparing those two pieces of code made the passing results easier to interpret: they had verified a calculation under a placement rule that was not used everywhere.
The project’s correction adds retries and a mapper-version check. That second change matters because two sides must agree on how an entry is placed for their summaries to be compatible. The project wrote and published the fix; my contribution was identifying the discrepancy and supplying evidence to reproduce it.
The README acknowledgment names Byungwoong Yoo, Independent Researcher. The fuller work record links the example and the later changes.
