Skip to content

Frequency Coding over Noisy Sampling

Bo-Yu Su, Hsin-Po Wang, Venkatesan Guruswami

cs.ITarXiv:2608.00539

Abstract

DNA molecules are so small that it might be practical to use their frequency vectors to encode messages. More precisely, a sender can inject MX copies of the string X = CATCATCAT into a pool and the receiver can recover MX by sequencing the pool. There are, however, two sources of uncertainty: (a) MX is usually too big to be counted exactly, but is estimated by sampling. (b) The DNA sequencer could be noisy; it may have difficulty distinguishing CATCATCAT from CATGATCAT. Recently, Tamir, Weinberger, and Guillén i Fàbregas clarified the amount of information the frequency vector can carry under (a). They showed that each string can carry about 4 R bits, where R is the average number of times each string is read. They also showed that 4 R bits can be achieved by a low-complexity uncoded scheme under the condition that there are at least R distinct strings. In this paper, we show that a low-complexity coded scheme can achieve the same 4 R bits unconditionally. We then generalize the scheme to handle sequencing noise, (b), and show that the noise penalizes the total number of bits by 2 W, together with a linear term due to the use of Fourier transforms in our proof. The former penalty 2 W is asymptotically the same as that obtained by Gerzon, Shomorony, and Weinberger; our scheme trades a small amount of rate for practical complexity.

Create a lesson