revisiting “cellular computing”
What does it really mean to say that biological systems process information?
The words “computing,” “information processing” and “circuit” (and related engineering language) are thrown around loosely in molecular biology and synthetic biology (including in earlier posts in this newsletter, unfortunately). Molecular biologists claim that processes like computation occur in various biological modalities, including gene expression, gene regulatory networks and signaling pathways, but there is no clear definition for most of the engineering language that they use. The lack of precise definitions for these words, which play important roles in the scientific language used in molecular biology and synthetic biology, has been bothering me for a while. This post will begin to explore some of the computing language used in biology from both engineering and biology perspectives. It’s just a few thoughts right now that I’m planning to explore more deeply in future posts. (But if you have anything to add, comment or message me!)
Computation
What does it mean to compute?
I like this definition from Wolfram MathWorld: “A computation is an operation that begins with some initial conditions and gives an output which follows from a definite set of rules.”
Or this one from Computer Science Stack Exchange: “a computation is just a mapping between some set 𝐴 to some set 𝐵” (emphasis in original text), where A is called the input and B is called the output. Although what we might add to the second definition, which the first already includes, is that the means by which the input is transformed into the output is well-defined and systematic.
Why do synthetic biologists say that cells compute?
Cells exhibit stereotyped responses to particular stimuli, which implies that they carry out processes that allow them to consistently convert specific inputs into specific output behavior.
Observations of these natural biological computations in single- and multicellular systems have inspired scientists to attempt building synthetic biological computers, an area of research that now goes interchangeably by the names of “synthetic biology” and “engineering biology” (not to be confused with the much broader field of “bioengineering”). Synthetic biologists focus mainly on computing within single cells and in populations of cells.
Information
What is information?
Merriam-Webster defines information as “knowledge gained from investigation, study, or instruction.”
The quantitative definition of information, given by Claude Shannon, the father of information theory, is how much a particular message changes what you know or expect about an event.
What is information processing?
Information processing is anything a system does with information, including storing, filtering, or transmitting it. While information processing fits the broad definition of computation above, engineers generally tend to think of computation as a specific case of information processing, in which the system applies algorithms (sets of logical rules or instructions) to transform inputs into outputs. (So because biological systems compute, it must also be true that (by the earlier definitions) they process information.)
What is biological information?
It is easy enough to quantify information at the level of DNA: at every base position, you can have 1 of 4 possible bases (A, T, C, G), making for 2 bits of information per base. (In fact, we can write to and read information from DNA just as we do with memory in computers.)
DNA is popularly known to encode information used to make proteins. The central dogma of molecular biology describes the flow of information within a cell, from DNA to RNA to protein (or rather, amino acid sequence). But the sequence of bases in a protein-coding region of DNA alone does not provide enough information to dictate the structure and function of a protein. Chemical modifications to the amino acid sequence and chaperone proteins which aid protein folding also provide some information about how certain proteins should be formed. Indeed, mammalian cells can produce fully formed proteins from a particular segment of DNA, but putting that same segment into bacteria does not lead to the same results (bacteria lack some of the capabilities that allow proteins to fold in specific ways).
Moreover, only a small number of DNA sequences out of the many that you can randomly generate actually carry some “meaning” in that they encode for specific, functional proteins. Biologists would probably consider a known protein-coding DNA sequence to be more informative than a random one, but might not be able to explain quantitatively why that is.
In general, it seems like biologists haven’t been too concerned with quantifying the flow of information in single cells and multicellular systems, or creating specific definitions for the plethora of ‘information words’ and engineering terminology that they use so frequently to describe these systems. At best, this ambiguity means that computing analogies to biological systems are simply left to be qualitative; at worst, it can lead to severe misunderstandings about how biological systems ‘work.’ At least one future post will get a little further into the word “information” as it pertains to molecular biology and synthetic biology, and maybe another will focus on the word “circuit” and other electrical engineering principles as they are used in synthetic biology.
This post is pretty stream-of-consciousness + certainly a work in progress. Comment or message if you have any thoughts to add!
Cover image courtesy Phys.org.

